Finding

Sign in with GitHub
← Publications

Finding · P284 · Author-curated

Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer on the Polish questions of 600 Belebele translation pairs from 0.4300 to 0.5933, a gain of 0.1633 (95% interval 0.1217 to 0.2050) that recovers 0.9074 of its English-minus-Polish accuracy gap of 0.1800.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Recovers the gap by

intervention
Translation assistChange applied to each Polish prompt.
evaluation set
Belebele translation pairsItems scored.
baseline
0.4300 proportionAccuracy on the Polish questions without the intervention.
treated
0.5933 proportionAccuracy on the Polish questions with the intervention.
gain
0.1633 proportionTreated minus baseline.
gain interval low
0.1217 proportionLower bound of the 95% paired bootstrap interval of the gain.
gain interval high
0.2050 proportionUpper bound of the 95% paired bootstrap interval of the gain.
gap
0.1800 proportionEnglish-minus-Polish accuracy gap of the model on these pairs.
recovery
0.9074 ratioGain divided by the gap.
continued pretraining tokens
500170752 tokensTokens of continued pretraining of the model at the checkpoint scored.

Experimental provenance

Method and evaluation protocol
Accuracy on the Polish questions with the English twin's full prompt and a blank line prepended, Polish scaffold kept, minus accuracy without it; recovery divides the gain by English minus Polish accuracy on the same pairs; length-normalised likelihood, torch MPS bf16, no hooks.
Dataset
600 Belebele and 600 MMLU translation pairs, all pairs without exclusions.Version: unspecified · Access: public
Reported results
accuracy English 0.6100, Polish 0.4300, Polish with assist 0.5933; gain 0.1633 [0.1217, 0.2050]; gap 0.1800; recovery 0.9074.
Uncertainty and replication
95% paired item bootstrap interval of the gain (2000 resamples, seed 20260703); no interval for the recovery ratio.
Limitations
All 600 pairs per benchmark: the audited exclusions are not applied in this stage. Letter-prior collapse in the models' multiple-choice predictions was not audited and may shift accuracy changes of the continued-pretraining arms. One scoring run per condition. The assist supplies the full English twin, so the recovery bounds how much of the gap depends on processing the Polish input; it does not measure Polish comprehension alone.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Recovers the gap by

For the model at continued_pretraining_tokens, the intervention raises the metric on the Polish questions from baseline to treated, a gain with the stated 95% interval that equals recovery times gap, the English-minus-Polish accuracy gap on the same pairs.

Key recovers_gap_by · version 803ee4c6-d174-4db8-be9d-f81f6313c7f1

Belebele translation pairs

600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.

Key belebele_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Translation assist

Prepending the English twin's full prompt and one blank line to the Polish prompt, which keeps its Polish answer scaffold.

Key translation_assist · version 489055f5-af11-44e3-b63d-c59d6d7cfef6

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584