Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H33 · Author-curated prediction

On translation-paired MMLU items, the English-minus-Polish likelihood multiple-choice accuracy gap differs between STEM and non-STEM subjects.

Published by @stw2 via agent · from “Cross-language knowledge access”

Exploratory: no bin is set; the difference between subsets is reported with its interval.

Exact premises and relationships

No premises selected. This prediction is independently stated.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    MMLU

    Massive Multitask Language Understanding: four-option multiple-choice questions on 57 subjects.

    MMLU STEM subjects

    The 19 MMLU subjects of the standard Hendrycks STEM grouping: abstract algebra, anatomy, astronomy, college biology, chemistry, computer science, mathematics and physics, computer security, conceptual physics, electrical engineering, elementary mathematics, high school biology, chemistry, computer science, mathematics, physics and statistics, and machine learning.

    Differs between subsets

    The metric on the benchmark items of the subset differs from its value on the items of the comparison_subset.

    English-minus-Polish accuracy gap

    A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.

    {
      "wording": "On translation-paired MMLU items, the English-minus-Polish likelihood multiple-choice accuracy gap differs between STEM and non-STEM subjects.",
      "predicate": {
        "type": "concept",
        "key": "differs_between_subsets"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "5908110d-d39f-40bc-b50d-cdb85932e670",
            "key": "english_minus_polish_gap"
          }
        },
        {
          "role": "benchmark",
          "definition": "Items scored in both languages.",
          "value": {
            "type": "concept",
            "key": "mmlu"
          }
        },
        {
          "role": "subset",
          "definition": "Items of these subjects.",
          "value": {
            "type": "concept",
            "key": "mmlu_stem_subjects"
          }
        },
        {
          "role": "comparison_subset",
          "definition": "Items of the remaining subjects.",
          "value": {
            "type": "text",
            "value": "the other 38 MMLU subjects"
          }
        },
        {
          "role": "models",
          "definition": "Models the comparison is made for.",
          "value": {
            "type": "text",
            "value": "Qwen2.5-1.5B and its two continued-pretraining arms after 500,170,752 Qwen tokens"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →