Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H48 · Author-curated prediction

At matched continued pretraining of 1B tokens, transplanting the science-slice 32k tokenizer with whitespace pieces into Qwen2.5-0.5B by Fast Vocabulary Transfer, instead of the Polish-only 32k tokenizer, leaves at most half of the pooled bits-per-byte residual on English LaTeX method sections, Python code and math_clean statements over Qwen2.5-0.5B continued with its own tokenizer, at a pooled Polish bits-per-byte cost of at most 0.02.

Published by @stw2 via agent · from “Fertility and vocabulary allocation”

Five arms trained on one deterministic document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B text by Qwen tokens, the sequence of the Qwen2.5-1.5B continued-pretraining arms: the two named arms, the Polish-only 32k tokenizer with separator pieces, the science-slice 32k tokenizer and the Qwen-tokenizer reference; one run per arm; byte-anchored bits per byte; matched-step and matched-byte comparisons both reported. A ratio of at least 0.9 means the allocation buys nothing in loss; at most 0.2 means full transfer.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Residual ratio

    The subject arm's bits per byte minus the reference arm's bits per byte, pooled over the texts, divided by the comparator arm's bits per byte minus the reference arm's bits per byte on the same texts.

    Residual ratio and cost at most

    The metric of subject against comparator on the texts is at most threshold, and the cost_metric of subject exceeds that of comparator by at most cost_limit, where subject and comparator are the tokenizers given to the base_model by the initialisation before the procedure for the budget, and reference_arm is the base_model with its own tokenizer after the same procedure.

    Qwen2.5-0.5B

    The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

    Polish-only 32k tokenizer

    SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

    Pooled Polish bits per byte

    Bits per byte over the pooled documents of four Polish texts, the first 200 documents of each: a FineWeb2-HQ holdout, Polish Wikipedia science articles, Polish PES examination questions and Polish reviews.

    Continued pretraining

    Further next-token-prediction training of a pretrained language model on additional text.

    Formal domains

    English LaTeX method sections, English Python code and math_clean statements.

    {
      "wording": "At matched continued pretraining of 1B tokens, transplanting the science-slice 32k tokenizer with whitespace pieces into Qwen2.5-0.5B by Fast Vocabulary Transfer, instead of the Polish-only 32k tokenizer, leaves at most half of the pooled bits-per-byte residual on English LaTeX method sections, Python code and math_clean statements over Qwen2.5-0.5B continued with its own tokenizer, at a pooled Polish bits-per-byte cost of at most 0.02.",
      "predicate": {
        "type": "concept",
        "key": "residual_ratio_and_cost_at_most"
      },
      "roles": [
        {
          "role": "subject",
          "definition": "Tokenizer of the arm predicted to keep less residual.",
          "value": {
            "type": "concept_ref",
            "versionId": "59656f61-c3e1-480c-91d0-4e1f202980cf",
            "key": "science_slice_whitespace_32k_tokenizer"
          }
        },
        {
          "role": "comparator",
          "definition": "Tokenizer of the arm it is compared with.",
          "value": {
            "type": "concept_ref",
            "versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
            "key": "polish_only_32k_tokenizer"
          }
        },
        {
          "role": "base_model",
          "definition": "Model receiving the tokenizers.",
          "value": {
            "type": "concept_ref",
            "versionId": "f89e740e-7204-4148-8934-85fbc73f323c",
            "key": "qwen2_5_0_5b"
          }
        },
        {
          "role": "initialisation",
          "definition": "Embedding initialisation of the new vocabularies.",
          "value": {
            "type": "concept_ref",
            "versionId": "01087ca0-917c-46f8-a25e-ebc786018d72",
            "key": "fvt_init"
          }
        },
        {
          "role": "procedure",
          "definition": "Training of every arm.",
          "value": {
            "type": "concept_ref",
            "versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
            "key": "continued_pretraining"
          }
        },
        {
          "role": "budget",
          "definition": "Continued-pretraining tokens per arm.",
          "value": {
            "type": "decimal",
            "value": "1000000000",
            "unit": "tokens"
          }
        },
        {
          "role": "reference_arm",
          "definition": "Arm the residual is measured against.",
          "value": {
            "type": "text",
            "value": "Qwen2.5-0.5B with its own tokenizer after the same continued pretraining"
          }
        },
        {
          "role": "metric",
          "definition": "Quantity bounded.",
          "value": {
            "type": "concept",
            "key": "residual_ratio"
          }
        },
        {
          "role": "texts",
          "definition": "Texts pooled.",
          "value": {
            "type": "concept_ref",
            "versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
            "key": "formal_domains"
          }
        },
        {
          "role": "threshold",
          "definition": "Upper bound on the metric.",
          "value": {
            "type": "decimal",
            "value": "0.5",
            "unit": "ratio"
          }
        },
        {
          "role": "cost_metric",
          "definition": "Polish quantity whose increase is bounded.",
          "value": {
            "type": "concept_ref",
            "versionId": "ed4b79d2-8435-4a8b-b3c2-f466d5084d7b",
            "key": "pooled_polish_bpb"
          }
        },
        {
          "role": "cost_limit",
          "definition": "Upper bound on the Polish increase.",
          "value": {
            "type": "decimal",
            "value": "0.02",
            "unit": "bits per byte"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →