Finding

Sign in with GitHub
← Publications

Finding · P246 · Author-curated

On 598 translation-paired Belebele items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B with APT4 by FVT after 500M tokens of Polish-heavy continued pretraining is 0.6104 in English and 0.4749 in Polish, an English-minus-Polish gap of 0.1355 (95% interval 0.0956 to 0.1753).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: English and Polish accuracy gap

metric
Likelihood multiple-choice accuracyQuantity measured in each language.
benchmark
Translation-paired Belebele itemsItem pairs scored.
subject
EnglishLanguage of the first value.
subject value
0.6104 accuracyAccuracy on the English items.
comparator
PolishLanguage of the second value.
comparator value
0.4749 accuracyAccuracy on the Polish items.
difference
0.1355 accuracyEnglish minus Polish accuracy.
interval low
0.0956 accuracyLower bound of the 95% interval of the difference.
interval high
0.1753 accuracyUpper bound of the 95% interval of the difference.
p value
0.0000000000866 probabilityExact two-sided sign test on discordant pairs.
training
500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per armTraining the continued-pretraining arms received.

Experimental provenance

Method and evaluation protocol
Each item pair is scored in English and Polish by the log-probability of each option letter after a scaffold in the item's language; accuracy uses length-normalised log-probabilities; the interval is the normal approximation for paired proportions and the p-value the exact sign test on discordant pairs.
Dataset
598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.Version: facebook/belebele@7899cdfa4e1e0d733fd77c848e2c273cb1d32be2 · Access: public
Reported results
English 0.6104, Polish 0.4749, gap 0.1355 [0.0956, 0.1753]; pairs correct in English only 120, in Polish only 39; sign test p 0.0000000000866.
Uncertainty and replication
Paired normal 95% interval over item pairs; exact two-sided sign test on discordant pairs.
Limitations
One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

English and Polish accuracy gap

subject_value and comparator_value are the setting model's metric on the benchmark item pairs in the subject and comparator languages; difference is subject_value minus comparator_value over the same pairs with its 95% interval from interval_low to interval_high; p_value is the exact two-sided sign test on the pairs answered correctly in only one language.

Key accuracy_gap · version e6ce775f-1927-4ab2-90fc-cdab7bcec624

Translation-paired Belebele items

598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.

Key belebele_paired_items · version e6ce775f-1927-4ab2-90fc-cdab7bcec624

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584