Finding

Sign in with GitHub
← Publications

Finding · P281 · Author-curated

Swapping the answer-scaffold language changes the likelihood-scored accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer after 500M tokens of continued pretraining by 0.0250 on Polish and -0.0483 on English questions of 600 Belebele translation pairs, and by 0.0383 and -0.0600 on 600 MMLU translation pairs; not every absolute change is below 0.03.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Changes by

intervention
Scaffold language swapChange applied to each prompt.
metric
Likelihood multiple-choice accuracyQuantity whose change is stated.
evaluation set 1
Belebele translation pairsFirst set of items.
evaluation set 2
MMLU translation pairsSecond set of items.
change set 1 polish
0.0250 proportionChange on the Polish questions of the first set.
change set 1 english
-0.0483 proportionChange on the English questions of the first set.
change set 2 polish
0.0383 proportionChange on the Polish questions of the second set.
change set 2 english
-0.0600 proportionChange on the English questions of the second set.
continued pretraining tokens
500170752 tokensTokens of continued pretraining of the model at the checkpoint scored.
threshold
0.03 proportionBound on the absolute change.
within threshold
falseWhether every absolute change is below threshold.

Experimental provenance

Method and evaluation protocol
Accuracy of each question language with the other language's answer scaffold minus accuracy with the matched scaffold, on the same items; length-normalised likelihood over four option letters, torch MPS bf16, no hooks.
Dataset
600 Belebele and 600 MMLU translation pairs, all pairs without exclusions.Version: unspecified · Access: public
Reported results
Belebele: Polish 0.0250, English -0.0483; MMLU: Polish 0.0383, English -0.0600; baseline accuracy Belebele English 0.6117, Polish 0.4767, MMLU English 0.4550, Polish 0.3617.
Uncertainty and replication
Point differences of accuracies on 600 items; no interval is computed for this contrast.
Limitations
All 600 pairs per benchmark: the audited exclusions are not applied in this stage. Letter-prior collapse in the models' multiple-choice predictions was not audited and may shift accuracy changes of the continued-pretraining arms. One scoring run per condition.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Changes by

The intervention changes the metric of the model, at continued_pretraining_tokens, by the stated amount on the Polish and on the English questions of each evaluation set; within_threshold states whether every absolute change is below threshold.

Key changes_by · version 6d5dcdcb-65c6-4efe-a0f8-67c7a935084b

Belebele translation pairs

600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.

Key belebele_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

MMLU translation pairs

600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.

Key mmlu_translation_pairs · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Scaffold language swap

Ending a prompt with the answer scaffold of the other language: 'Odpowiedź (litera):' after an English question, 'Answer (letter):' after a Polish question.

Key scaffold_language_swap · version cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584