MMLU
Massive Multitask Language Understanding: four-option multiple-choice questions on 57 subjects.
Hypothesis
Sign in with GitHubHypothesis · H33 · Author-curated prediction
Exploratory: no bin is set; the difference between subsets is reported with its interval.
No premises selected. This prediction is independently stated.
Loading research…
Massive Multitask Language Understanding: four-option multiple-choice questions on 57 subjects.
The 19 MMLU subjects of the standard Hendrycks STEM grouping: abstract algebra, anatomy, astronomy, college biology, chemistry, computer science, mathematics and physics, computer security, conceptual physics, electrical engineering, elementary mathematics, high school biology, chemistry, computer science, mathematics, physics and statistics, and machine learning.
The metric on the benchmark items of the subset differs from its value on the items of the comparison_subset.
A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.
{
"wording": "On translation-paired MMLU items, the English-minus-Polish likelihood multiple-choice accuracy gap differs between STEM and non-STEM subjects.",
"predicate": {
"type": "concept",
"key": "differs_between_subsets"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept_ref",
"versionId": "5908110d-d39f-40bc-b50d-cdb85932e670",
"key": "english_minus_polish_gap"
}
},
{
"role": "benchmark",
"definition": "Items scored in both languages.",
"value": {
"type": "concept",
"key": "mmlu"
}
},
{
"role": "subset",
"definition": "Items of these subjects.",
"value": {
"type": "concept",
"key": "mmlu_stem_subjects"
}
},
{
"role": "comparison_subset",
"definition": "Items of the remaining subjects.",
"value": {
"type": "text",
"value": "the other 38 MMLU subjects"
}
},
{
"role": "models",
"definition": "Models the comparison is made for.",
"value": {
"type": "text",
"value": "Qwen2.5-1.5B and its two continued-pretraining arms after 500,170,752 Qwen tokens"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →