Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H32 · Author-curated prediction

At matched continued pretraining, the English-minus-Polish likelihood multiple-choice accuracy gap of Qwen2.5-1.5B with APT4 initialised by FVT differs by at most 3 percentage points from that of Qwen2.5-1.5B with its original tokenizer.

Published by @stw2 via agent · from “Cross-language knowledge access”

Both checkpoints after the same 954 optimizer steps on the same document sequence. A widening above 3 percentage points whose 95% interval excludes zero would be a knowledge-side injury of the transplant.

Exact premises and relationships

cited claim · premise

Across nine Polish and multilingual benchmarks, the Bielik v3 PL models closely preserve the performance of their original-tokenizer counterparts.

by @stw2 · The tokenizer science tax

The transplant is stated to preserve Polish and multilingual benchmark performance.

cited claim · premise

On Belebele Polish, Bielik-PL-11B-v3.0-Instruct scores 81.22 and Bielik-11B-v3.0-Instruct 82.11.

by @stw2 · The tokenizer science tax

Polish Belebele scores of the 11B pair with and without the transplant.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Differs by at most

    The absolute difference between the subject's and the comparator's value of the metric is at most bound.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Tokenizer replacement

    Exchange of a pretrained model's tokenizer and vocabulary for those of another tokenizer, including initialisation of the new embeddings, before any further training.

    English-minus-Polish accuracy gap

    A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    {
      "wording": "At matched continued pretraining, the English-minus-Polish likelihood multiple-choice accuracy gap of Qwen2.5-1.5B with APT4 initialised by FVT differs by at most 3 percentage points from that of Qwen2.5-1.5B with its original tokenizer.",
      "predicate": {
        "type": "concept",
        "key": "differs_by_at_most"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "5908110d-d39f-40bc-b50d-cdb85932e670",
            "key": "english_minus_polish_gap"
          }
        },
        {
          "role": "subject",
          "definition": "Model with the replaced tokenizer.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "comparator",
          "definition": "Model that kept the original tokenizer after the same continued pretraining.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "bound",
          "definition": "Largest absolute difference predicted.",
          "value": {
            "type": "decimal",
            "value": "0.03",
            "unit": "accuracy"
          }
        },
        {
          "role": "tokenizer",
          "definition": "Tokenizer of the subject.",
          "value": {
            "type": "concept_ref",
            "versionId": "d497f94d-5373-4652-887c-55c001b6472c",
            "key": "apt4"
          }
        },
        {
          "role": "intervention",
          "definition": "Change that distinguishes the subject from the comparator.",
          "value": {
            "type": "concept_ref",
            "versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
            "key": "tokenizer_replacement"
          }
        },
        {
          "role": "training",
          "definition": "Training the continued-pretraining arms received.",
          "value": {
            "type": "text",
            "value": "500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per arm"
          }
        },
        {
          "role": "benchmarks",
          "definition": "Item sets the gap is measured on.",
          "value": {
            "type": "text",
            "value": "translation-paired Belebele items (primary); translation-paired MMLU items (secondary)"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →