Finding

Sign in with GitHub
← Publications

Finding · P282 · Author-curated

Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B on the Polish questions of 600 Belebele translation pairs from 0.5483 to 0.6367, a gain of 0.0883 (95% interval 0.0483 to 0.1317) that recovers 0.4454 of its English-minus-Polish accuracy gap of 0.1983.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Recovers the gap by

model
Qwen2.5-1.5BModel scored.
intervention
Translation assistChange applied to each Polish prompt.
evaluation set
Belebele translation pairsItems scored.
baseline
0.5483 proportionAccuracy on the Polish questions without the intervention.
treated
0.6367 proportionAccuracy on the Polish questions with the intervention.
gain
0.0883 proportionTreated minus baseline.
gain interval low
0.0483 proportionLower bound of the 95% paired bootstrap interval of the gain.
gain interval high
0.1317 proportionUpper bound of the 95% paired bootstrap interval of the gain.
gap
0.1983 proportionEnglish-minus-Polish accuracy gap of the model on these pairs.
recovery
0.4454 ratioGain divided by the gap.
continued pretraining tokens
0 tokensTokens of continued pretraining of the model at the checkpoint scored.

Experimental provenance

Method and evaluation protocol
Accuracy on the Polish questions with the English twin's full prompt and a blank line prepended, Polish scaffold kept, minus accuracy without it; recovery divides the gain by English minus Polish accuracy on the same pairs; length-normalised likelihood, torch MPS bf16, no hooks.
Dataset
600 Belebele and 600 MMLU translation pairs, all pairs without exclusions.Version: unspecified · Access: public
Reported results
accuracy English 0.7467, Polish 0.5483, Polish with assist 0.6367; gain 0.0883 [0.0483, 0.1317]; gap 0.1983; recovery 0.4454.
Uncertainty and replication
95% paired item bootstrap interval of the gain (2000 resamples, seed 20260703); no interval for the recovery ratio.
Limitations
All 600 pairs per benchmark: the audited exclusions are not applied in this stage. Letter-prior collapse in the models' multiple-choice predictions was not audited and may shift accuracy changes of the continued-pretraining arms. One scoring run per condition. The assist supplies the full English twin, so the recovery bounds how much of the gap depends on processing the Polish input; it does not measure Polish comprehension alone.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Recovers the gap by

For the model at continued_pretraining_tokens, the intervention raises the metric on the Polish questions from baseline to treated, a gain with the stated 95% interval that equals recovery times gap, the English-minus-Polish accuracy gap on the same pairs.

Key recovers_gap_by · version 803ee4c6-d174-4db8-be9d-f81f6313c7f1

Belebele translation pairs

600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.

Key belebele_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Translation assist

Prepending the English twin's full prompt and one blank line to the Polish prompt, which keeps its Polish answer scaffold.

Key translation_assist · version 489055f5-af11-44e3-b63d-c59d6d7cfef6

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Exact references

refines

e6ce775f-1927-4ab2-90fc-cdab7bcec624

Share of this gap recovered when the English twin is in the prompt.