Finding

Sign in with GitHub
← Publications

Finding · P279 · Author-curated

Swapping the answer-scaffold language changes the likelihood-scored accuracy of Qwen2.5-1.5B by 0.0233 on Polish and -0.0017 on English questions of 600 Belebele translation pairs, and by 0.0150 and 0.0033 on 600 MMLU translation pairs; every absolute change is below 0.03.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Changes by

model
Qwen2.5-1.5BModel scored.
intervention
Scaffold language swapChange applied to each prompt.
metric
Likelihood multiple-choice accuracyQuantity whose change is stated.
evaluation set 1
Belebele translation pairsFirst set of items.
evaluation set 2
MMLU translation pairsSecond set of items.
change set 1 polish
0.0233 proportionChange on the Polish questions of the first set.
change set 1 english
-0.0017 proportionChange on the English questions of the first set.
change set 2 polish
0.0150 proportionChange on the Polish questions of the second set.
change set 2 english
0.0033 proportionChange on the English questions of the second set.
continued pretraining tokens
0 tokensTokens of continued pretraining of the model at the checkpoint scored.
threshold
0.03 proportionBound on the absolute change.
within threshold
trueWhether every absolute change is below threshold.

Experimental provenance

Method and evaluation protocol
Accuracy of each question language with the other language's answer scaffold minus accuracy with the matched scaffold, on the same items; length-normalised likelihood over four option letters, torch MPS bf16, no hooks.
Dataset
600 Belebele and 600 MMLU translation pairs, all pairs without exclusions.Version: unspecified · Access: public
Reported results
Belebele: Polish 0.0233, English -0.0017; MMLU: Polish 0.0150, English 0.0033; baseline accuracy Belebele English 0.7467, Polish 0.5483, MMLU English 0.5533, Polish 0.3400.
Uncertainty and replication
Point differences of accuracies on 600 items; no interval is computed for this contrast.
Limitations
All 600 pairs per benchmark: the audited exclusions are not applied in this stage. Letter-prior collapse in the models' multiple-choice predictions was not audited and may shift accuracy changes of the continued-pretraining arms. One scoring run per condition.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Changes by

The intervention changes the metric of the model, at continued_pretraining_tokens, by the stated amount on the Polish and on the English questions of each evaluation set; within_threshold states whether every absolute change is below threshold.

Key changes_by · version 6d5dcdcb-65c6-4efe-a0f8-67c7a935084b

Belebele translation pairs

600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.

Key belebele_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

MMLU translation pairs

600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.

Key mmlu_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Scaffold language swap

Ending a prompt with the answer scaffold of the other language: 'Odpowiedź (litera):' after an English question, 'Answer (letter):' after a Polish question.

Key scaffold_language_swap · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Exact references

refines

e6ce775f-1927-4ab2-90fc-cdab7bcec624

Whether the gap tracks question language rather than scaffold format.

refines

22e9d053-dd03-4cb2-ba2d-bea9e71a1260

Whether the gap tracks question language rather than scaffold format.