Ratio exceeds a threshold
The metric's mean over subject_domains divided by its mean over comparator_domains exceeds threshold.
Hypothesis
Sign in with GitHubHypothesis · H20 · Author-curated prediction
Bits per byte of APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings, and of the base model, on the first 200 documents of nine evaluation domains, before any training.
cited claim · premise
FOCUS represents each token of the target vocabulary as a sparse linear combination of tokens from the original vocabulary, selected by semantic similarity in an auxiliary embedding space.The informed initialisation whose advantage over random is measured.
cited claim · premise
Random initialisation assigns randomly sampled vectors to new tokens, requiring the model to relearn embeddings from scratch and often resulting in slow convergence.The floor the advantage is measured against.
cited claim · premise
The choice of FOCUS is supported by prior experiments on earlier Bielik v3 models that evaluated multiple embedding initialisation strategies, in which FOCUS consistently showed the best empirical performance; on Bielik 1.5B v3 it gave the lowest training loss after 4B tokens of continued pretraining and leading results on the Open Polish LLM Leaderboard.The initialisation choice rests on aggregate results at 1.5B.
Loading research…
The metric's mean over subject_domains divided by its mean over comparator_domains exceeds threshold.
Bits per byte of the random-initialised transplant minus bits per byte of the FOCUS-initialised transplant on the same documents.
English arXiv abstracts, Polish Wikipedia science articles, Polish PES examination questions, the Polish FineWeb2-HQ holdout, the English SlimPajama holdout and Polish reviews.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and a new 32,000-row input embedding matrix, tied to the output head, filled by an embedding initialisation.
English LaTeX method sections, English Python code and math_clean statements.
{
"wording": "At time zero, the mean init-quality spread of APT4 transplants of Qwen2.5-1.5B over formal domains exceeds 1.25 times its mean over prose domains.",
"predicate": {
"type": "concept",
"key": "ratio_exceeds"
},
"roles": [
{
"role": "metric",
"definition": "Quantity averaged over each group of domains.",
"value": {
"type": "concept",
"key": "init_quality_spread"
}
},
{
"role": "model",
"definition": "Models measured.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "qwen_apt4_transplant"
}
},
{
"role": "subject_domains",
"definition": "Domains whose mean is the numerator.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "formal_domains"
}
},
{
"role": "comparator_domains",
"definition": "Domains whose mean is the denominator.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "prose_domains"
}
},
{
"role": "setting",
"definition": "Training state of the models.",
"value": {
"type": "text",
"value": "time zero: no training after embedding initialisation"
}
},
{
"role": "threshold",
"definition": "Value the ratio is predicted to exceed.",
"value": {
"type": "decimal",
"value": "1.25",
"unit": "ratio"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Embedding initialisation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →