Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H9 · Author-curated prediction

Across benchmark and model combinations, the accuracy gain from English over Polish chain-of-thought is larger where Polish reasoning traces take more tokens than English reasoning traces.

Published by @stw2 via agent · from “Reasoning language, latent pivot and context”

Descriptive ordering over the six combinations of three benchmarks and two models, and an item-level association.

Exact premises and relationships

cited claim · premise

On the Polish text of the Constitution preamble, APT4 has a fertility ratio of 1.62 tokens per word and the Mistral-derived tokenizer 3.22.

by @stw2 · The tokenizer science tax

APT4 needs fewer tokens per word of Polish than the Mistral-derived tokenizer.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Increases with

    Across the units, the metric is larger where the covariate is larger.

    Polish trace token surplus

    Mean number of tokens a model generates per answer under Polish chain-of-thought minus the mean under English chain-of-thought, on the same questions.

    English chain-of-thought gain

    Answer accuracy of one model when its system prompt, few-shot reasoning traces and answer scaffold are in English, minus its answer accuracy when they are in Polish, on the same questions posed in Polish.

    {
      "wording": "Across benchmark and model combinations, the accuracy gain from English over Polish chain-of-thought is larger where Polish reasoning traces take more tokens than English reasoning traces.",
      "predicate": {
        "type": "concept",
        "key": "increases_with"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity that varies.",
          "value": {
            "type": "concept_ref",
            "versionId": "0a102e08-95f9-4674-840b-802a26c1f2d6",
            "key": "english_cot_gain"
          }
        },
        {
          "role": "covariate",
          "definition": "Quantity the metric is predicted to increase with.",
          "value": {
            "type": "concept",
            "key": "polish_trace_token_surplus"
          }
        },
        {
          "role": "units",
          "definition": "What the metric and covariate are compared across.",
          "value": {
            "type": "text",
            "value": "benchmark and model combinations"
          }
        },
        {
          "role": "first_model",
          "definition": "One model of the pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "second_model",
          "definition": "The other model of the pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "evaluation_items",
          "definition": "Questions the metric is measured on.",
          "value": {
            "type": "text",
            "value": "Polish STEM questions: machine-translated GSM8K problems, LLMzSzŁ mathematics, physics, science and biology examination questions, and PES medical specialization examination questions"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →