Mean exceeds
The metric's mean over subject_domains exceeds its mean over comparator_domains.
Hypothesis
Sign in with GitHubHypothesis · H21 · Author-curated prediction
Bits per byte of APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings, and of the base model, on the first 200 documents of nine evaluation domains, before any training.
cited claim · premise
FOCUS represents each token of the target vocabulary as a sparse linear combination of tokens from the original vocabulary, selected by semantic similarity in an auxiliary embedding space.FOCUS, whose advantage over FVT is compared across domains.
cited claim · premise
Frequency-based Vocabulary Transfer initialises token embeddings by aggregating representations of their constituent subword units, guided by frequency statistics.FVT, the cheaper initialisation compared with.
cited claim · premise
The choice of FOCUS is supported by prior experiments on earlier Bielik v3 models that evaluated multiple embedding initialisation strategies, in which FOCUS consistently showed the best empirical performance; on Bielik 1.5B v3 it gave the lowest training loss after 4B tokens of continued pretraining and leading results on the Open Polish LLM Leaderboard.The choice of FOCUS over the other methods rests on aggregate results.
Loading research…
The metric's mean over subject_domains exceeds its mean over comparator_domains.
Bits per byte of the FVT-initialised transplant minus bits per byte of the FOCUS-initialised transplant on the same documents.
English arXiv abstracts, Polish Wikipedia science articles, Polish PES examination questions, the Polish FineWeb2-HQ holdout, the English SlimPajama holdout and Polish reviews.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and a new 32,000-row input embedding matrix, tied to the output head, filled by an embedding initialisation.
English LaTeX method sections, English Python code and math_clean statements.
{
"wording": "At time zero, the mean FVT−FOCUS gap of APT4 transplants of Qwen2.5-1.5B over formal domains exceeds its mean over prose domains.",
"predicate": {
"type": "concept",
"key": "mean_exceeds"
},
"roles": [
{
"role": "metric",
"definition": "Quantity averaged over each group of domains.",
"value": {
"type": "concept",
"key": "focus_fvt_gap"
}
},
{
"role": "model",
"definition": "Models measured.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "qwen_apt4_transplant"
}
},
{
"role": "subject_domains",
"definition": "Domains whose mean is predicted to be larger.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "formal_domains"
}
},
{
"role": "comparator_domains",
"definition": "Domains whose mean is predicted to be smaller.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "prose_domains"
}
},
{
"role": "setting",
"definition": "Training state of the models.",
"value": {
"type": "text",
"value": "time zero: no training after embedding initialisation"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Embedding initialisation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →