Cited claim

Sign in with GitHub
← Publications

Cited claim · P26 · Author-curated

Vocabulary adaptation of the Bielik v3 PL models combines FOCUS-based embedding initialisation with continued pretraining on 4B tokens that updates only the input embedding layer, the language modelling head and four boundary transformer layers, followed by 16B tokens with all parameters unfrozen, to mitigate catastrophic forgetting.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Follows a procedure

subject
Bielik v3 PL modelsModels adapted.
stage
Vocabulary adaptationTraining stage.
initialisation
FOCUSEmbedding initialisation.
stage one tokens
4000000000 tokensContinued-pretraining tokens in stage one.
stage one trainable
input embedding layer; language modelling head; four boundary transformer layers (two lowest, two highest)Parameters updated in stage one.
stage two tokens
16000000000 tokensContinued-pretraining tokens in stage two.
stage two trainable
all parametersParameters updated in stage two.
purpose
mitigate catastrophic forgetting during vocabulary adaptationStated purpose of the procedure.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Section 7, p. 15

    To mitigate catastrophic forgetting during vocabulary adaptation, we combined FOCUS-based embedding initialization with a two-stage continued pretraining pipeline (4B tokens with partial freezing, followed by 16B tokens of full adaptation)
  2. Section 4.1.1, p. 4

    Continued pretraining is performed on 4B tokens, while most of the model parameters remain frozen. Only the following components are updated: • The input embedding layer, • The language modeling head (lm_head), • Four boundary transformer layers (two lowest and two highest layers).
  3. Section 4.1.2, p. 4

    After initial stabilization, all model parameters are unfrozen. The model then undergoes continued pretraining on an additional 16B tokens.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Follows a procedure

The subject's training stage runs the listed steps in order, for the stated purpose.

Key follows_procedure · version 4aeee3b3-fb89-4ae6-9e40-9668ffa62e84

FOCUS

Fast Overlapping Token Combinations Using Sparsemax: an embedding initialisation for a replaced tokenizer's vocabulary.

Key focus_init · version 5a078e71-84ff-4039-9e32-7999a5f679f5

Vocabulary adaptation

Adaptation of a pretrained model to a replaced tokenizer's vocabulary.

Key vocabulary_adaptation · version 95d5752d-3660-4d7c-97d5-f23b933cf433

Bielik v3 PL models

The 11B and 7B Bielik v3 models with the APT4 tokenizer.

Key bielik_v3_pl_models · version 3da77265-eb29-46f1-90e2-ccc03ac36917