Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H47 · Author-curated prediction

On natively Polish LLMzSzŁ STEM and PES examination items and their machine translations into English, the permutation-marginalised likelihood multiple-choice accuracy of Qwen2.5-1.5B is higher in English than in Polish.

Published by @stw2 via agent · from “Cross-language knowledge access”

Qwen2.5-1.5B; translation from Polish into English, so translation effects fall on the English side; the gap's size is compared, without a bin, with the same model's gap on translation-paired Belebele and MMLU items under the same scoring; per-item round-trip translation quality is a covariate.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Higher in a language

    The subject's metric on the benchmark items in language is higher than on the same items in comparison_language.

    Qwen2.5-1.5B

    The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

    {
      "wording": "On natively Polish LLMzSzŁ STEM and PES examination items and their machine translations into English, the permutation-marginalised likelihood multiple-choice accuracy of Qwen2.5-1.5B is higher in English than in Polish.",
      "predicate": {
        "type": "concept_ref",
        "versionId": "50041e38-5826-4b54-b39b-fdff21892c61",
        "key": "higher_in_language"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "3e921b79-6353-4934-9191-9793faaa652e",
            "key": "permutation_marginalised_mcq_accuracy"
          }
        },
        {
          "role": "subject",
          "definition": "Model scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
            "key": "qwen2_5_1_5b"
          }
        },
        {
          "role": "benchmark",
          "definition": "Items scored in both languages.",
          "value": {
            "type": "concept",
            "key": "native_polish_exam_items_english_mt"
          }
        },
        {
          "role": "language",
          "definition": "Language predicted to score higher.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "english"
          }
        },
        {
          "role": "comparison_language",
          "definition": "Language of the same items predicted to score lower.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "polish"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →