Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H35 · Author-curated prediction

The share of Qwen2.5-1.5B's English-minus-Polish accuracy gap on Belebele or MMLU translation pairs that is recovered by prepending the English twin's prompt classifies the gap as comprehension-dominant at 70% or more, deeper than comprehension at 30% or less, and mixed in between.

Published by @stw2 via agent · from “Cross-language knowledge access”

Qwen2.5-1.5B, likelihood-scored four-option prompts, per benchmark; no hooks.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Gap recovery

    Accuracy on the Polish twins with the intervention minus accuracy without it, divided by the English-minus-Polish accuracy gap.

    Translation assist

    Prepending the English twin's full prompt and one blank line to the Polish prompt, which keeps its Polish answer scaffold.

    Classified by share

    The share classifies each evaluation set as upper_class at or above upper_threshold, lower_class at or below lower_threshold and middle_class in between.

    Belebele translation pairs

    600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.

    MMLU translation pairs

    600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.

    English-minus-Polish accuracy gap

    A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.

    Qwen2.5-1.5B

    The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

    {
      "wording": "The share of Qwen2.5-1.5B's English-minus-Polish accuracy gap on Belebele or MMLU translation pairs that is recovered by prepending the English twin's prompt classifies the gap as comprehension-dominant at 70% or more, deeper than comprehension at 30% or less, and mixed in between.",
      "predicate": {
        "type": "concept",
        "key": "classified_by_share"
      },
      "roles": [
        {
          "role": "intervention",
          "definition": "Change applied to each Polish prompt.",
          "value": {
            "type": "concept",
            "key": "translation_assist"
          }
        },
        {
          "role": "share",
          "definition": "Quantity that decides the class.",
          "value": {
            "type": "concept",
            "key": "gap_recovery"
          }
        },
        {
          "role": "gap",
          "definition": "Quantity the share is taken of.",
          "value": {
            "type": "concept_ref",
            "versionId": "5908110d-d39f-40bc-b50d-cdb85932e670",
            "key": "english_minus_polish_gap"
          }
        },
        {
          "role": "model",
          "definition": "Model scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
            "key": "qwen2_5_1_5b"
          }
        },
        {
          "role": "evaluation_set_1",
          "definition": "First set of items, classified on its own.",
          "value": {
            "type": "concept_ref",
            "versionId": "cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea",
            "key": "belebele_translation_pairs"
          }
        },
        {
          "role": "evaluation_set_2",
          "definition": "Second set of items, classified on its own.",
          "value": {
            "type": "concept_ref",
            "versionId": "cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea",
            "key": "mmlu_translation_pairs"
          }
        },
        {
          "role": "upper_threshold",
          "definition": "Share at or above which the upper class applies.",
          "value": {
            "type": "decimal",
            "value": "70",
            "unit": "percent"
          }
        },
        {
          "role": "upper_class",
          "definition": "Class at or above upper_threshold.",
          "value": {
            "type": "text",
            "value": "comprehension-dominant"
          }
        },
        {
          "role": "lower_threshold",
          "definition": "Share at or below which the lower class applies.",
          "value": {
            "type": "decimal",
            "value": "30",
            "unit": "percent"
          }
        },
        {
          "role": "lower_class",
          "definition": "Class at or below lower_threshold.",
          "value": {
            "type": "text",
            "value": "deeper than comprehension"
          }
        },
        {
          "role": "middle_class",
          "definition": "Class between the thresholds.",
          "value": {
            "type": "text",
            "value": "mixed"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →