Cited claim

Sign in with GitHub
← Publications

Cited claim · P24 · Author-curated

The choice of FOCUS is supported by prior experiments on earlier Bielik v3 models that evaluated multiple embedding initialisation strategies, in which FOCUS consistently showed the best empirical performance; on Bielik 1.5B v3 it gave the lowest training loss after 4B tokens of continued pretraining and leading results on the Open Polish LLM Leaderboard.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Supported by evidence

choice
FOCUSThe chosen method.
methods considered
Random Initialization; Frequency-based Vocabulary Transfer (FVT); Linear Transformation (aX + b); WECHSEL; FOCUS; MATT; OFA; RAMENMethods listed as considered.
evidence origin
prior experiments on earlier Bielik v3 modelsWhere the evidence comes from.
overall result
consistently the best empirical performanceSummary of the evidence.
evidence model
Bielik 1.5B v3Model the named results were obtained on.
loss result
lowest training lossTraining-loss result.
loss budget
4000000000 tokensContinued-pretraining tokens at which the loss was compared.
benchmark result
leading results on the Open Polish LLM LeaderboardBenchmark result.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Section 4, p. 3

    Our choice of FOCUS is supported by prior experimental results on earlier Bielik v3 models Ociepa et al. [2025b], where multiple embedding initialization strategies were systematically evaluated.
  2. Section 4, pp. 3-4, list of methods considered: Random Initialization, Frequency-based Vocabulary Transfer (FVT), Linear Transformation (aX + b), WECHSEL, FOCUS, MATT, OFA, RAMEN

  3. Section 4, p. 4

    Among these approaches, FOCUS consistently demonstrated the best empirical performance. In particular, experiments conducted on the Bielik 1.5B v3 model showed the lowest training loss after 4B tokens of continued pretraining, as well as leading results on the Open Polish LLM Leaderboard Wróbel et al. [2024], Ociepa et al. [2025c].

Author’s note

The evidence is reported in Ociepa et al. [2025b]; this paper gives no measurement of it.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Bielik 1.5B v3

The 1.5B-parameter Bielik v3 model.

Key bielik_1_5b_v3 · version f5ada7c4-30c8-405f-8741-cd6fcfb58530

Supported by evidence

The choice is supported by the evidence stated in the evidence and result roles.

Key supported_by_evidence · version f5ada7c4-30c8-405f-8741-cd6fcfb58530

FOCUS

Fast Overlapping Token Combinations Using Sparsemax: an embedding initialisation for a replaced tokenizer's vocabulary.

Key focus_init · version 5a078e71-84ff-4039-9e32-7999a5f679f5