Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H44 · Author-curated prediction

On English LaTeX method sections, Python code and math_clean statements, at least 60% of the excess bits of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer over Qwen2.5-1.5B with its own tokenizer, after the same 500M tokens of continued pretraining, fall in the lowest predictive-entropy tercile.

Published by @stw2 via agent · from “The missing control arm”

Checkpoints after 500,170,752 Qwen tokens of continued pretraining, one run per arm; byte-anchored evaluation windows; the partition must agree under both projection rules; the Polish texts and the continued-pretraining erosion of the original-tokenizer arm are negative controls. Terciles are checked against the untouched model.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Share at least

    At least share_threshold of the measure of the subject over the comparator, summed over the texts, falls on the byte_class, under every projection rule.

    Lowest predictive-entropy tercile

    Bytes of the original-tokenizer arm's tokens whose next-token predictive entropy under that arm lies in the lowest third of that arm's tokens on the text.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Excess bits on bytes

    For each byte of a text, the bits a subject model assigns to it minus the bits a comparator model assigns to it, each model's next-token negative log-likelihood in bits being projected from each token onto that token's bytes by a projection rule.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    Formal domains

    English LaTeX method sections, English Python code and math_clean statements.

    {
      "wording": "On English LaTeX method sections, Python code and math_clean statements, at least 60% of the excess bits of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer over Qwen2.5-1.5B with its own tokenizer, after the same 500M tokens of continued pretraining, fall in the lowest predictive-entropy tercile.",
      "predicate": {
        "type": "concept",
        "key": "share_at_least"
      },
      "roles": [
        {
          "role": "measure",
          "definition": "Quantity partitioned.",
          "value": {
            "type": "concept_ref",
            "versionId": "10d42787-ae0d-4bf2-ad51-7d292d298f14",
            "key": "excess_bits"
          }
        },
        {
          "role": "subject",
          "definition": "Model whose bits are the minuend.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "comparator",
          "definition": "Model whose bits are the subtrahend.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "continued_pretraining_tokens",
          "definition": "Tokens of continued pretraining of both arms at the checkpoints named.",
          "value": {
            "type": "decimal",
            "value": "500170752",
            "unit": "tokens"
          }
        },
        {
          "role": "texts",
          "definition": "Texts scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
            "key": "formal_domains"
          }
        },
        {
          "role": "byte_class",
          "definition": "Bytes predicted to carry the excess.",
          "value": {
            "type": "concept",
            "key": "lowest_entropy_tercile"
          }
        },
        {
          "role": "share_threshold",
          "definition": "Lower bound on the share of the summed measure.",
          "value": {
            "type": "decimal",
            "value": "0.6",
            "unit": "ratio"
          }
        },
        {
          "role": "projection_rules",
          "definition": "Rules projecting a token's bits onto its bytes.",
          "value": {
            "type": "text",
            "value": "uniform over the token's bytes; all bits on the token's final byte"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →