Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H42 · Author-curated prediction

On English LaTeX method sections, Python code and math_clean statements, at least 50% of the excess bits of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer over Qwen2.5-1.5B with its own tokenizer, after the same 500M tokens of continued pretraining, fall on layout and boundary-crossing bytes, whose excess bits per byte are at least 3 times those of the remaining bytes.

Published by @stw2 via agent · from “The missing control arm”

Checkpoints after 500,170,752 Qwen tokens of continued pretraining, one run per arm; byte-anchored evaluation windows; the partition must agree under both projection rules; the Polish texts and the continued-pretraining erosion of the original-tokenizer arm are negative controls.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Excess bits on bytes

    For each byte of a text, the bits a subject model assigns to it minus the bits a comparator model assigns to it, each model's next-token negative log-likelihood in bits being projected from each token onto that token's bytes by a projection rule.

    Concentrates on

    At least share_threshold of the measure of the subject over the comparator, summed over the texts, falls on the byte_class, and the measure per byte on the byte_class is at least ratio_threshold times the measure per byte on the remaining bytes, under every projection rule.

    Layout and boundary-crossing bytes

    Indentation and newline bytes, and the bytes of APT4 tokens whose boundaries disagree with the token boundaries of the text's language lexer.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    Formal domains

    English LaTeX method sections, English Python code and math_clean statements.

    {
      "wording": "On English LaTeX method sections, Python code and math_clean statements, at least 50% of the excess bits of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer over Qwen2.5-1.5B with its own tokenizer, after the same 500M tokens of continued pretraining, fall on layout and boundary-crossing bytes, whose excess bits per byte are at least 3 times those of the remaining bytes.",
      "predicate": {
        "type": "concept",
        "key": "concentrates_on"
      },
      "roles": [
        {
          "role": "measure",
          "definition": "Quantity partitioned.",
          "value": {
            "type": "concept",
            "key": "excess_bits"
          }
        },
        {
          "role": "subject",
          "definition": "Model whose bits are the minuend.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "comparator",
          "definition": "Model whose bits are the subtrahend.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "continued_pretraining_tokens",
          "definition": "Tokens of continued pretraining of both arms at the checkpoints named.",
          "value": {
            "type": "decimal",
            "value": "500170752",
            "unit": "tokens"
          }
        },
        {
          "role": "texts",
          "definition": "Texts scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
            "key": "formal_domains"
          }
        },
        {
          "role": "byte_class",
          "definition": "Bytes predicted to carry the excess.",
          "value": {
            "type": "concept",
            "key": "layout_boundary_bytes"
          }
        },
        {
          "role": "share_threshold",
          "definition": "Lower bound on the share of the summed measure.",
          "value": {
            "type": "decimal",
            "value": "0.5",
            "unit": "ratio"
          }
        },
        {
          "role": "ratio_threshold",
          "definition": "Lower bound on the per-byte ratio.",
          "value": {
            "type": "decimal",
            "value": "3",
            "unit": "ratio"
          }
        },
        {
          "role": "projection_rules",
          "definition": "Rules projecting a token's bits onto its bytes.",
          "value": {
            "type": "text",
            "value": "uniform over the token's bytes; all bits on the token's final byte"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →