Finding

Sign in with GitHub
← Publications

Finding · P280 · Author-curated

Swapping the answer-scaffold language changes the likelihood-scored accuracy of Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer by 0.0167 on Polish and 0.0017 on English questions of 600 Belebele translation pairs, and by 0.0183 and -0.0233 on 600 MMLU translation pairs; every absolute change is below 0.03.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Changes by

intervention
Scaffold language swapChange applied to each prompt.
metric
Likelihood multiple-choice accuracyQuantity whose change is stated.
evaluation set 1
Belebele translation pairsFirst set of items.
evaluation set 2
MMLU translation pairsSecond set of items.
change set 1 polish
0.0167 proportionChange on the Polish questions of the first set.
change set 1 english
0.0017 proportionChange on the English questions of the first set.
change set 2 polish
0.0183 proportionChange on the Polish questions of the second set.
change set 2 english
-0.0233 proportionChange on the English questions of the second set.
continued pretraining tokens
500170752 tokensTokens of continued pretraining of the model at the checkpoint scored.
threshold
0.03 proportionBound on the absolute change.
within threshold
trueWhether every absolute change is below threshold.

Experimental provenance

Method and evaluation protocol
Accuracy of each question language with the other language's answer scaffold minus accuracy with the matched scaffold, on the same items; length-normalised likelihood over four option letters, torch MPS bf16, no hooks.
Dataset
600 Belebele and 600 MMLU translation pairs, all pairs without exclusions.Version: unspecified · Access: public
Reported results
Belebele: Polish 0.0167, English 0.0017; MMLU: Polish 0.0183, English -0.0233; baseline accuracy Belebele English 0.6100, Polish 0.4300, MMLU English 0.4500, Polish 0.3417.
Uncertainty and replication
Point differences of accuracies on 600 items; no interval is computed for this contrast.
Limitations
All 600 pairs per benchmark: the audited exclusions are not applied in this stage. Letter-prior collapse in the models' multiple-choice predictions was not audited and may shift accuracy changes of the continued-pretraining arms. One scoring run per condition.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Changes by

The intervention changes the metric of the model, at continued_pretraining_tokens, by the stated amount on the Polish and on the English questions of each evaluation set; within_threshold states whether every absolute change is below threshold.

Key changes_by · version 6d5dcdcb-65c6-4efe-a0f8-67c7a935084b

Belebele translation pairs

600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.

Key belebele_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

MMLU translation pairs

600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.

Key mmlu_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Scaffold language swap

Ending a prompt with the answer scaffold of the other language: 'Odpowiedź (litera):' after an English question, 'Answer (letter):' after a Polish question.

Key scaffold_language_swap · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584