Finding

Sign in with GitHub
← Publications

Finding · P287 · Author-curated

Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer after 500M tokens of continued pretraining on the Polish questions of 600 MMLU translation pairs from 0.3617 to 0.4200, a gain of 0.0583 (95% interval 0.0266 to 0.0917) that recovers 0.6250 of its English-minus-Polish accuracy gap of 0.0933.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Recovers the gap by

intervention
Translation assistChange applied to each Polish prompt.
evaluation set
MMLU translation pairsItems scored.
baseline
0.3617 proportionAccuracy on the Polish questions without the intervention.
treated
0.4200 proportionAccuracy on the Polish questions with the intervention.
gain
0.0583 proportionTreated minus baseline.
gain interval low
0.0266 proportionLower bound of the 95% paired bootstrap interval of the gain.
gain interval high
0.0917 proportionUpper bound of the 95% paired bootstrap interval of the gain.
gap
0.0933 proportionEnglish-minus-Polish accuracy gap of the model on these pairs.
recovery
0.6250 ratioGain divided by the gap.
continued pretraining tokens
500170752 tokensTokens of continued pretraining of the model at the checkpoint scored.

Experimental provenance

Method and evaluation protocol
Accuracy on the Polish questions with the English twin's full prompt and a blank line prepended, Polish scaffold kept, minus accuracy without it; recovery divides the gain by English minus Polish accuracy on the same pairs; length-normalised likelihood, torch MPS bf16, no hooks.
Dataset
600 Belebele and 600 MMLU translation pairs, all pairs without exclusions.Version: unspecified · Access: public
Reported results
accuracy English 0.4550, Polish 0.3617, Polish with assist 0.4200; gain 0.0583 [0.0266, 0.0917]; gap 0.0933; recovery 0.6250.
Uncertainty and replication
95% paired item bootstrap interval of the gain (2000 resamples, seed 20260703); no interval for the recovery ratio.
Limitations
All 600 pairs per benchmark: the audited exclusions are not applied in this stage. Letter-prior collapse in the models' multiple-choice predictions was not audited and may shift accuracy changes of the continued-pretraining arms. One scoring run per condition. The assist supplies the full English twin, so the recovery bounds how much of the gap depends on processing the Polish input; it does not measure Polish comprehension alone.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Recovers the gap by

For the model at continued_pretraining_tokens, the intervention raises the metric on the Polish questions from baseline to treated, a gain with the stated 95% interval that equals recovery times gap, the English-minus-Polish accuracy gap on the same pairs.

Key recovers_gap_by · version 803ee4c6-d174-4db8-be9d-f81f6313c7f1

MMLU translation pairs

600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.

Key mmlu_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Translation assist

Prepending the English twin's full prompt and one blank line to the Polish prompt, which keeps its Polish answer scaffold.

Key translation_assist · version 489055f5-af11-44e3-b63d-c59d6d7cfef6

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584