Finding

Sign in with GitHub
← Publications

Finding · P247 · Author-curated

On 589 translation-paired MMLU items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.5569 in English and 0.3430 in Polish, an English-minus-Polish gap of 0.2139 (95% interval 0.1672 to 0.2606).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: English and Polish accuracy gap

metric
Likelihood multiple-choice accuracyQuantity measured in each language.
benchmark
Translation-paired MMLU itemsItem pairs scored.
setting
Qwen2.5-1.5BModel scored.
subject
EnglishLanguage of the first value.
subject value
0.5569 accuracyAccuracy on the English items.
comparator
PolishLanguage of the second value.
comparator value
0.3430 accuracyAccuracy on the Polish items.
difference
0.2139 accuracyEnglish minus Polish accuracy.
interval low
0.1672 accuracyLower bound of the 95% interval of the difference.
interval high
0.2606 accuracyUpper bound of the 95% interval of the difference.
p value
0.00000000000000000838 probabilityExact two-sided sign test on discordant pairs.

Experimental provenance

Method and evaluation protocol
Each item pair is scored in English and Polish by the log-probability of each option letter after a scaffold in the item's language; accuracy uses length-normalised log-probabilities; the interval is the normal approximation for paired proportions and the p-value the exact sign test on discordant pairs.
Dataset
589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.Version: cais/mmlu@c30699e8356da336a370243923dbaf21066bb9fe; openGPT-X/mmlux@7fb62cf05c1ce9ea550dbfd2ea6c145204344882 · Access: public
Reported results
English 0.5569, Polish 0.3430, gap 0.2139 [0.1672, 0.2606]; pairs correct in English only 175, in Polish only 49; sign test p 0.00000000000000000838.
Uncertainty and replication
Paired normal 95% interval over item pairs; exact two-sided sign test on discordant pairs.
Limitations
One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Translation-paired MMLU items

589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.

Key mmlu_paired_items · version 22e9d053-dd03-4cb2-ba2d-bea9e71a1260

English and Polish accuracy gap

subject_value and comparator_value are the setting model's metric on the benchmark item pairs in the subject and comparator languages; difference is subject_value minus comparator_value over the same pairs with its 95% interval from interval_low to interval_high; p_value is the exact two-sided sign test on the pairs answered correctly in only one language.

Key accuracy_gap · version e6ce775f-1927-4ab2-90fc-cdab7bcec624

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Exact references

related

cbcebd2f-77d5-48ae-9534-159d89911958

Multilingual examination scores of the Bielik 11B pair.