Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H31 · Author-curated prediction

Continued pretraining of Qwen2.5-1.5B on 500M tokens of an 80/20 Polish-English mix narrows its English-minus-Polish likelihood multiple-choice accuracy gap on translation-paired items.

Published by @stw2 via agent · from “Cross-language knowledge access”

Belebele is the primary benchmark and MMLU, with machine-translated Polish items, the secondary. A narrowing counts as improved Polish access only if Polish accuracy rises with a 95% interval excluding zero; a narrowing carried by falling English accuracy with flat Polish accuracy is erosion.

Exact premises and relationships

cited claim · premise

Across nine Polish and multilingual benchmarks, the Bielik v3 PL models closely preserve the performance of their original-tokenizer counterparts.

by @stw2 · The tokenizer science tax

Performance on Polish and multilingual benchmarks is stated to be preserved after vocabulary adaptation with continued pretraining.

cited claim · premise

On INCLUDE-base-44, averaged over 20 European languages, Bielik-PL-11B-v3.0-Instruct scores 53.92 and Bielik-11B-v3.0-Instruct 64.8.

by @stw2 · The tokenizer science tax

A multilingual benchmark average falls after the transplant with continued pretraining.

finding · premise

Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens lowers bits per byte on the Polish FineWeb2-HQ holdout from 1.2043 to 0.9345.

by @stw2 · The tokenizer science tax

Continued pretraining with the original tokenizer lowers Polish bits per byte, so the arm scored here did learn Polish text.

finding · premise

Untouched Qwen2.5-1.5B has 0.0524 bits per byte on GSM8K problems, the lowest of 10 texts scored; the next lowest is 0.3417 on Python code.

by @stw2 · The tokenizer science tax

A likely memorised English benchmark text has the base model's lowest bits per byte.

finding · premise

Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens raises bits per byte on GSM8K problems from 0.0524 to 0.4966.

by @stw2 · The tokenizer science tax

Continued pretraining raises bits per byte on that memorised text, a decay that can lower English accuracy without any change in access.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Narrows

    The subject's value of the metric is smaller than the comparator's.

    English-minus-Polish accuracy gap

    A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.

    Qwen2.5-1.5B

    The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

    Continued pretraining

    Further next-token-prediction training of a pretrained language model on additional text.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    {
      "wording": "Continued pretraining of Qwen2.5-1.5B on 500M tokens of an 80/20 Polish-English mix narrows its English-minus-Polish likelihood multiple-choice accuracy gap on translation-paired items.",
      "predicate": {
        "type": "concept",
        "key": "narrows"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "english_minus_polish_gap"
          }
        },
        {
          "role": "subject",
          "definition": "Model after the intervention.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "comparator",
          "definition": "Model before the intervention.",
          "value": {
            "type": "concept_ref",
            "versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
            "key": "qwen2_5_1_5b"
          }
        },
        {
          "role": "intervention",
          "definition": "Training applied to the comparator to obtain the subject.",
          "value": {
            "type": "concept_ref",
            "versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
            "key": "continued_pretraining"
          }
        },
        {
          "role": "training",
          "definition": "Training the continued-pretraining arms received.",
          "value": {
            "type": "text",
            "value": "500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per arm"
          }
        },
        {
          "role": "benchmarks",
          "definition": "Item sets the gap is measured on.",
          "value": {
            "type": "text",
            "value": "translation-paired Belebele items (primary); translation-paired MMLU items (secondary)"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →