Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H43 · Author-curated prediction

On English LaTeX method sections, Python code and math_clean statements, the excess bits per byte of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer over Qwen2.5-1.5B with its own tokenizer, after the same 500M tokens of continued pretraining, are at least 2 times as high at copy positions as at other positions and rise with the number of APT4 pieces of the copied unit.

Published by @stw2 via agent · from “The missing control arm”

Checkpoints after 500,170,752 Qwen tokens of continued pretraining, one run per arm; byte-anchored evaluation windows; the partition must agree under both projection rules; the Polish texts and the continued-pretraining erosion of the original-tokenizer arm are negative controls.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Copy positions

    Bytes inside an n-gram of at least 8 bytes that occurs earlier in the same document.

    At least a ratio and rising

    The measure per byte of the subject over the comparator on the texts is at least ratio_threshold times as high on the byte_class as on the remaining bytes, and on the byte_class it increases with the covariate, under every projection rule.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Excess bits on bytes

    For each byte of a text, the bits a subject model assigns to it minus the bits a comparator model assigns to it, each model's next-token negative log-likelihood in bits being projected from each token onto that token's bytes by a projection rule.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    Formal domains

    English LaTeX method sections, English Python code and math_clean statements.

    {
      "wording": "On English LaTeX method sections, Python code and math_clean statements, the excess bits per byte of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer over Qwen2.5-1.5B with its own tokenizer, after the same 500M tokens of continued pretraining, are at least 2 times as high at copy positions as at other positions and rise with the number of APT4 pieces of the copied unit.",
      "predicate": {
        "type": "concept",
        "key": "ratio_at_least_and_rising"
      },
      "roles": [
        {
          "role": "measure",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "10d42787-ae0d-4bf2-ad51-7d292d298f14",
            "key": "excess_bits"
          }
        },
        {
          "role": "subject",
          "definition": "Model whose bits are the minuend.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "comparator",
          "definition": "Model whose bits are the subtrahend.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "continued_pretraining_tokens",
          "definition": "Tokens of continued pretraining of both arms at the checkpoints named.",
          "value": {
            "type": "decimal",
            "value": "500170752",
            "unit": "tokens"
          }
        },
        {
          "role": "texts",
          "definition": "Texts scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
            "key": "formal_domains"
          }
        },
        {
          "role": "byte_class",
          "definition": "Bytes predicted to carry more excess.",
          "value": {
            "type": "concept",
            "key": "copy_positions"
          }
        },
        {
          "role": "covariate",
          "definition": "Quantity the measure is predicted to rise with.",
          "value": {
            "type": "text",
            "value": "number of APT4 pieces of the copied unit"
          }
        },
        {
          "role": "ratio_threshold",
          "definition": "Lower bound on the per-byte ratio.",
          "value": {
            "type": "decimal",
            "value": "2",
            "unit": "ratio"
          }
        },
        {
          "role": "projection_rules",
          "definition": "Rules projecting a token's bits onto its bytes.",
          "value": {
            "type": "text",
            "value": "uniform over the token's bytes; all bits on the token's final byte"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →