Finding

Sign in with GitHub
← Publications

Finding · P270 · Author-curated

On 589 translation-paired MMLU items, Qwen2.5-1.5B's English likelihood multiple-choice accuracy is 0.3964 on the 111 STEM pairs and 0.5941 on the 478 non-STEM pairs.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Likelihood multiple-choice accuracyQuantity compared.
benchmark
Translation-paired MMLU itemsItem pairs scored.
language
EnglishLanguage of the items.
setting
Qwen2.5-1.5BModel scored.
subject
MMLU STEM subjectsFirst subset.
subject value
0.3964 accuracyAccuracy on the STEM pairs.
comparator
MMLU pairs of the other 38 subjectsSecond subset.
comparator value
0.5941 accuracyAccuracy on the non-STEM pairs.

Experimental provenance

Method and evaluation protocol
Length-normalised accuracy of the base model's English rows of scores_base.jsonl, exclusions applied, split by the STEM tag.
Dataset
589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.Version: cais/mmlu@c30699e8356da336a370243923dbaf21066bb9fe; openGPT-X/mmlux@7fb62cf05c1ce9ea550dbfd2ea6c145204344882 · Access: public
Reported results
STEM 0.3964, non-STEM 0.5941.
Uncertainty and replication
Point values without intervals.
Limitations
Computed in the port from the committed score file; the original report states the same values. One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Translation-paired MMLU items

589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.

Key mmlu_paired_items · version 22e9d053-dd03-4cb2-ba2d-bea9e71a1260

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

MMLU STEM subjects

The 19 MMLU subjects of the standard Hendrycks STEM grouping: abstract algebra, anatomy, astronomy, college biology, chemistry, computer science, mathematics and physics, computer security, conceptual physics, electrical engineering, elementary mathematics, high school biology, chemistry, computer science, mathematics, physics and statistics, and machine learning.

Key mmlu_stem_subjects · version 50390697-7a2b-44cd-bf3c-ffd92d092fa7

Exact references

related

7960d8ed-895f-4679-82fd-033f5e8e77f6

The STEM gap is measured on items with lower English accuracy.