Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H8 · Author-curated prediction

On Polish STEM questions, the accuracy gain from English over Polish chain-of-thought is larger for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct.

Published by @stw2 via agent · from “Reasoning language, latent pivot and context”

Two 11B instruction-tuned Bielik v3 models that differ in tokenizer; questions posed in Polish; the language of the reasoning is set by the prompt.

Exact premises and relationships

cited claim · premise

On GSM8K, Bielik-PL-11B-v3.0-Instruct scores 80.97 and Bielik-11B-v3.0-Instruct 85.60.

by @stw2 · The tokenizer science tax

The APT4 model scores lower on GSM8K than its original-tokenizer counterpart.

cited claim · premise

English-language capabilities of the Bielik v3 PL models remain largely intact.

by @stw2 · The tokenizer science tax

The statement the GSM8K values bear on.

cited claim · premise

On the Polish Medical Leaderboard, Bielik-PL-11B-v3.0-Instruct scores 48.42 percent and Bielik-11B-v3.0-Instruct 50.21 percent.

by @stw2 · The tokenizer science tax

The APT4 model scores lower on the PES-based benchmark than its counterpart.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    English chain-of-thought gain

    Answer accuracy of one model when its system prompt, few-shot reasoning traces and answer scaffold are in English, minus its answer accuracy when they are in Polish, on the same questions posed in Polish.

    Larger for the subject

    The metric is larger for the subject than for the comparator on the evaluation items, in the language and setting stated.

    {
      "wording": "On Polish STEM questions, the accuracy gain from English over Polish chain-of-thought is larger for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct.",
      "predicate": {
        "type": "concept",
        "key": "larger_for_subject"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "english_cot_gain"
          }
        },
        {
          "role": "subject",
          "definition": "Model predicted to have the larger value.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "comparator",
          "definition": "Model compared with.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "evaluation_items",
          "definition": "Questions the metric is measured on.",
          "value": {
            "type": "text",
            "value": "Polish STEM questions: machine-translated GSM8K problems, LLMzSzŁ mathematics, physics, science and biology examination questions, and PES medical specialization examination questions"
          }
        },
        {
          "role": "language",
          "definition": "Language of the questions.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "polish"
          }
        },
        {
          "role": "setting",
          "definition": "Evaluation setting.",
          "value": {
            "type": "text",
            "value": "few-shot prompts; greedy decoding"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →