Changes by less than
The intervention changes the metric by less than threshold in absolute value for every model, evaluation set and question language named.
Hypothesis
Sign in with GitHubHypothesis · H34 · Author-curated prediction
Qwen2.5-1.5B and its two continued-pretraining arms at 500M tokens, likelihood-scored four-option prompts; no hooks.
finding · premise
On 598 translation-paired Belebele items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.7458 in English and 0.5485 in Polish, an English-minus-Polish gap of 0.1973 (95% interval 0.1558 to 0.2389).The English-minus-Polish accuracy gap of the base model on these translation pairs.
finding · premise
On 589 translation-paired MMLU items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.5569 in English and 0.3430 in Polish, an English-minus-Polish gap of 0.2139 (95% interval 0.1672 to 0.2606).The English-minus-Polish accuracy gap of the base model on these translation pairs.
Loading research…
The intervention changes the metric by less than threshold in absolute value for every model, evaluation set and question language named.
600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.
Ending a prompt with the answer scaffold of the other language: 'Odpowiedź (litera):' after an English question, 'Answer (letter):' after a Polish question.
600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.
Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
The English language.
The Polish language.
{
"wording": "Answering Belebele and MMLU translation pairs with the answer scaffold in the other language changes the likelihood-scored accuracy of Qwen2.5-1.5B and of its two continued-pretraining arms at 500M tokens by less than 3 percentage points, for English and for Polish questions.",
"predicate": {
"type": "concept",
"key": "changes_less_than"
},
"roles": [
{
"role": "intervention",
"definition": "Change applied to each prompt.",
"value": {
"type": "concept",
"key": "scaffold_language_swap"
}
},
{
"role": "metric",
"definition": "Quantity whose change is bounded.",
"value": {
"type": "concept_ref",
"versionId": "50041e38-5826-4b54-b39b-fdff21892c61",
"key": "paired_mcq_accuracy"
}
},
{
"role": "model_1",
"definition": "First model scored.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "model_2",
"definition": "Second model scored, at its 500M-token checkpoint.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "model_3",
"definition": "Third model scored, at its 500M-token checkpoint.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "arm_continued_pretraining_tokens",
"definition": "Tokens of continued pretraining of the second and third models at the checkpoint scored.",
"value": {
"type": "decimal",
"value": "500170752",
"unit": "tokens"
}
},
{
"role": "evaluation_set_1",
"definition": "First set of items.",
"value": {
"type": "concept",
"key": "belebele_translation_pairs"
}
},
{
"role": "evaluation_set_2",
"definition": "Second set of items.",
"value": {
"type": "concept",
"key": "mmlu_translation_pairs"
}
},
{
"role": "question_language_1",
"definition": "First question language.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "english"
}
},
{
"role": "question_language_2",
"definition": "Second question language.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "polish"
}
},
{
"role": "threshold",
"definition": "Bound on the absolute change.",
"value": {
"type": "decimal",
"value": "3",
"unit": "percentage points"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →