Init damage
Bits per byte of a transplant on a domain minus bits per byte of the base model with its own tokenizer on the same documents.
Hypothesis
Sign in with GitHubHypothesis · H18 · Author-curated prediction
Bits per byte of APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings, and of the base model, on the first 200 documents of nine evaluation domains, before any training.
cited claim · premise
Random initialisation assigns randomly sampled vectors to new tokens, requiring the model to relearn embeddings from scratch and often resulting in slow convergence.Random initialisation, one of the two initialisations predicted to damage formal domains most.
cited claim · premise
Frequency-based Vocabulary Transfer initialises token embeddings by aggregating representations of their constituent subword units, guided by frequency statistics.FVT, the other initialisation.
cited claim · premise
The choice of FOCUS is supported by prior experiments on earlier Bielik v3 models that evaluated multiple embedding initialisation strategies, in which FOCUS consistently showed the best empirical performance; on Bielik 1.5B v3 it gave the lowest training loss after 4B tokens of continued pretraining and leading results on the Open Polish LLM Leaderboard.The initialisation choice rests on aggregate results at 1.5B, without per-domain comparison.
Loading research…
Bits per byte of a transplant on a domain minus bits per byte of the base model with its own tokenizer on the same documents.
English arXiv abstracts, Polish Wikipedia science articles, Polish PES examination questions, the Polish FineWeb2-HQ holdout, the English SlimPajama holdout and Polish reviews.
English LaTeX method sections, English Python code and math_clean statements.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and a new 32,000-row input embedding matrix, tied to the output head, filled by an embedding initialisation.
For initialisation_a and initialisation_b, the metric's mean over subject_domains divided by its mean over comparator_domains exceeds threshold.
An embedding initialisation for a replaced tokenizer's vocabulary that samples new vectors at random.
An embedding initialisation for a replaced tokenizer's vocabulary.
{
"wording": "At time zero, the mean init damage of APT4 transplants of Qwen2.5-1.5B over formal domains exceeds 1.15 times its mean over prose domains for both random and FVT initialisation.",
"predicate": {
"type": "concept",
"key": "ratio_exceeds_for_each"
},
"roles": [
{
"role": "metric",
"definition": "Quantity averaged over each group of domains.",
"value": {
"type": "concept",
"key": "init_damage"
}
},
{
"role": "model",
"definition": "Models measured.",
"value": {
"type": "concept",
"key": "qwen_apt4_transplant"
}
},
{
"role": "initialisation_a",
"definition": "First initialisation for which the ratio is predicted to exceed the threshold.",
"value": {
"type": "concept_ref",
"versionId": "cae5a2ea-f055-4e01-bb4b-86bb0b6c0fd3",
"key": "random_init"
}
},
{
"role": "initialisation_b",
"definition": "Second initialisation for which the ratio is predicted to exceed the threshold.",
"value": {
"type": "concept_ref",
"versionId": "01087ca0-917c-46f8-a25e-ebc786018d72",
"key": "fvt_init"
}
},
{
"role": "subject_domains",
"definition": "Domains whose mean is the numerator.",
"value": {
"type": "concept",
"key": "formal_domains"
}
},
{
"role": "comparator_domains",
"definition": "Domains whose mean is the denominator.",
"value": {
"type": "concept",
"key": "prose_domains"
}
},
{
"role": "setting",
"definition": "Training state of the models.",
"value": {
"type": "text",
"value": "time zero: no training after embedding initialisation"
}
},
{
"role": "threshold",
"definition": "Value each ratio is predicted to exceed.",
"value": {
"type": "decimal",
"value": "1.15",
"unit": "ratio"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Embedding initialisation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →