Within of
The subject's metric differs from the comparator's by at most threshold, and the 95% interval of the difference includes zero.
Hypothesis
Sign in with GitHubHypothesis · H39 · Author-curated prediction
The original-tokenizer arm at 500M tokens and Qwen2.5-1.5B on the base model's forward-discordant pairs; not adjudicated unless the validity gate passes.
finding · premise
After 500M tokens of Polish-heavy continued pretraining with its original tokenizer, Qwen2.5-1.5B's Polish likelihood multiple-choice accuracy rises on neither translation-paired Belebele items (difference -0.1171, 95% interval -0.1605 to -0.0702) nor translation-paired MMLU items (difference 0.0000, 95% interval -0.0458 to 0.0424).The behavioural result of continued pretraining on the same arm.
Loading research…
The subject's metric differs from the comparator's by at most threshold, and the 95% interval of the difference includes zero.
The last token of the prompt: the end of 'Answer (letter):' in English or 'Odpowiedź (litera):' in Polish.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
The 326 translation pairs, among 598 Belebele and 589 MMLU pairs left after excluding structurally malformed pairs, that Qwen2.5-1.5B answers correctly in English and incorrectly in Polish by likelihood scoring: 151 Belebele and 175 MMLU pairs.
Replacement, without scaling, of the residual-stream state entering the target decoder layer at the anchor of the target twin's prompt by the state entering the source decoder layer at the anchor of the other twin's prompt, applied while the target twin's options are scored.
Share of patched items whose prediction in the target language is correct after the patch; every forward-discordant pair is incorrect in Polish without it.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
{
"wording": "At the confirmed layer pairs, the scaffold-final residual patch has a flip rate on Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer within 10 percentage points of the flip rate on Qwen2.5-1.5B, with a 95% interval of the difference that includes zero.",
"predicate": {
"type": "concept",
"key": "within_of"
},
"roles": [
{
"role": "subject",
"definition": "Model whose metric is compared, at its 500M-token checkpoint.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "subject_continued_pretraining_tokens",
"definition": "Tokens of continued pretraining of the subject at the checkpoint patched.",
"value": {
"type": "decimal",
"value": "500170752",
"unit": "tokens"
}
},
{
"role": "comparator",
"definition": "Model compared with.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept_ref",
"versionId": "cd49225b-c235-425c-83bc-e24271555357",
"key": "flip_rate"
}
},
{
"role": "intervention",
"definition": "Change applied.",
"value": {
"type": "concept_ref",
"versionId": "cd49225b-c235-425c-83bc-e24271555357",
"key": "residual_patch"
}
},
{
"role": "anchor",
"definition": "Token position patched.",
"value": {
"type": "concept_ref",
"versionId": "cd49225b-c235-425c-83bc-e24271555357",
"key": "scaffold_final_anchor"
}
},
{
"role": "evaluation_set",
"definition": "Items patched.",
"value": {
"type": "concept_ref",
"versionId": "cd49225b-c235-425c-83bc-e24271555357",
"key": "forward_discordant_pairs"
}
},
{
"role": "threshold",
"definition": "Bound on the absolute difference.",
"value": {
"type": "decimal",
"value": "10",
"unit": "percentage points"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →