Finding

Sign in with GitHub
← Publications

Finding · P267 · Author-curated

In an exploratory split of 589 translation-paired MMLU items, the English-minus-Polish accuracy gap of Qwen2.5-1.5B is 0.0541 (95% interval -0.0453 to 0.1534) on the 111 STEM pairs and 0.2510 (0.1989 to 0.3032) on the 478 non-STEM pairs, a difference of -0.1970.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Gap on a subset and the rest

metric
English-minus-Polish accuracy gapQuantity compared.
benchmark
Translation-paired MMLU itemsItem pairs scored.
setting
Qwen2.5-1.5BModel scored.
subset
MMLU STEM subjectsSubjects whose pairs form the subset.
subset value
0.0541 accuracyGap on the STEM pairs.
subset interval low
-0.0453 accuracyLower bound of the 95% interval of subset_value.
subset interval high
0.1534 accuracyUpper bound of the 95% interval of subset_value.
rest value
0.2510 accuracyGap on the pairs of the other 38 subjects.
rest interval low
0.1989 accuracyLower bound of the 95% interval of rest_value.
rest interval high
0.3032 accuracyUpper bound of the 95% interval of rest_value.
difference
-0.1970 accuracySTEM gap minus non-STEM gap.

Experimental provenance

Method and evaluation protocol
Gap and paired normal 95% interval computed separately on the STEM and non-STEM pairs; the difference is the STEM gap minus the non-STEM gap.
Dataset
589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.Version: cais/mmlu@c30699e8356da336a370243923dbaf21066bb9fe; openGPT-X/mmlux@7fb62cf05c1ce9ea550dbfd2ea6c145204344882 · Access: public
Reported results
STEM 0.0541 [-0.0453, 0.1534]; non-STEM 0.2510 [0.1989, 0.3032]; difference -0.1970.
Uncertainty and replication
Per-subset paired normal 95% intervals; the design asked for an interval on the difference, and the analysis reports none.
Limitations
Exploratory, with no bin. Difficulty compression: the base model's English accuracy is 0.3964 on STEM and 0.5941 on non-STEM pairs, so STEM gaps sit closer to the 0.25 floor. One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Gap on a subset and the rest

subset_value is the setting model's metric on the benchmark item pairs in subset and rest_value on the remaining pairs of the benchmark, each with its 95% interval; difference is subset_value minus rest_value, without an interval.

Key subset_gap_comparison · version 7960d8ed-895f-4679-82fd-033f5e8e77f6

Translation-paired MMLU items

589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.

Key mmlu_paired_items · version 22e9d053-dd03-4cb2-ba2d-bea9e71a1260

English-minus-Polish accuracy gap

A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.

Key english_minus_polish_gap · version 5908110d-d39f-40bc-b50d-cdb85932e670

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

MMLU STEM subjects

The 19 MMLU subjects of the standard Hendrycks STEM grouping: abstract algebra, anatomy, astronomy, college biology, chemistry, computer science, mathematics and physics, computer security, conceptual physics, electrical engineering, elementary mathematics, high school biology, chemistry, computer science, mathematics, physics and statistics, and machine learning.

Key mmlu_stem_subjects · version 50390697-7a2b-44cd-bf3c-ffd92d092fa7