Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H30 · Author-curated prediction

Qwen2.5-1.5B's likelihood multiple-choice accuracy on translation-paired Belebele items is higher in English than in Polish.

Published by @stw2 via agent · from “Cross-language knowledge access”

Belebele test items with the same passage, question and options in English and Polish, each pair scored by the same model; comparison per item pair.

Exact premises and relationships

cited claim · premise

On Belebele Polish, Bielik-PL-11B-v3.0-Instruct scores 81.22 and Bielik-11B-v3.0-Instruct 82.11.

by @stw2 · The tokenizer science tax

Belebele measures Polish reading comprehension of the Bielik v3 models.

cited claim · premise

On Belebele, averaged over 28 European language variants, Bielik-PL-11B-v3.0-Instruct scores 77.41 and Bielik-11B-v3.0-Instruct 82.98.

by @stw2 · The tokenizer science tax

Belebele scores across European languages before and after the transplant.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Higher in a language

    The subject's metric on the benchmark items in language is higher than on the same items in comparison_language.

    Likelihood multiple-choice accuracy

    Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

    Belebele

    Multilingual multiple-choice reading comprehension benchmark on FLORES-200 passages.

    Qwen2.5-1.5B

    The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

    {
      "wording": "Qwen2.5-1.5B's likelihood multiple-choice accuracy on translation-paired Belebele items is higher in English than in Polish.",
      "predicate": {
        "type": "concept",
        "key": "higher_in_language"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "paired_mcq_accuracy"
          }
        },
        {
          "role": "subject",
          "definition": "Model scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
            "key": "qwen2_5_1_5b"
          }
        },
        {
          "role": "benchmark",
          "definition": "Items scored in both languages.",
          "value": {
            "type": "concept_ref",
            "versionId": "e0184bf2-dccb-469b-85be-cefff666bd50",
            "key": "belebele"
          }
        },
        {
          "role": "language",
          "definition": "Language predicted to score higher.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "english"
          }
        },
        {
          "role": "comparison_language",
          "definition": "Language of the same items predicted to score lower.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "polish"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →