Finding

Sign in with GitHub
← Publications

Finding · P245 · Author-curated

On 598 translation-paired Belebele items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B after 500M tokens of Polish-heavy continued pretraining with its original tokenizer is 0.6087 in English and 0.4314 in Polish, an English-minus-Polish gap of 0.1773 (95% interval 0.1378 to 0.2168).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: English and Polish accuracy gap

metric
Likelihood multiple-choice accuracyQuantity measured in each language.
benchmark
Translation-paired Belebele itemsItem pairs scored.
subject
EnglishLanguage of the first value.
subject value
0.6087 accuracyAccuracy on the English items.
comparator
PolishLanguage of the second value.
comparator value
0.4314 accuracyAccuracy on the Polish items.
difference
0.1773 accuracyEnglish minus Polish accuracy.
interval low
0.1378 accuracyLower bound of the 95% interval of the difference.
interval high
0.2168 accuracyUpper bound of the 95% interval of the difference.
p value
0.0000000000000000150 probabilityExact two-sided sign test on discordant pairs.
training
500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per armTraining the continued-pretraining arms received.

Experimental provenance

Method and evaluation protocol
Each item pair is scored in English and Polish by the log-probability of each option letter after a scaffold in the item's language; accuracy uses length-normalised log-probabilities; the interval is the normal approximation for paired proportions and the p-value the exact sign test on discordant pairs.
Dataset
598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.Version: facebook/belebele@7899cdfa4e1e0d733fd77c848e2c273cb1d32be2 · Access: public
Reported results
English 0.6087, Polish 0.4314, gap 0.1773 [0.1378, 0.2168]; pairs correct in English only 135, in Polish only 29; sign test p 0.0000000000000000150.
Uncertainty and replication
Paired normal 95% interval over item pairs; exact two-sided sign test on discordant pairs.
Limitations
One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

English and Polish accuracy gap

subject_value and comparator_value are the setting model's metric on the benchmark item pairs in the subject and comparator languages; difference is subject_value minus comparator_value over the same pairs with its 95% interval from interval_low to interval_high; p_value is the exact two-sided sign test on the pairs answered correctly in only one language.

Key accuracy_gap · version e6ce775f-1927-4ab2-90fc-cdab7bcec624

Translation-paired Belebele items

598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.

Key belebele_paired_items · version e6ce775f-1927-4ab2-90fc-cdab7bcec624

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584