Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H50 · Author-curated prediction

On English Python code and math_clean statements, word reconstruction depth mediates at least 30% of the per-token surprisal penalty of the APT4 model over its original-tokenizer counterpart, in the Bielik 11B v3 pair and in the Qwen2.5-1.5B continued-pretraining pair.

Published by @stw2 via agent · from “Fertility and vocabulary allocation”

Bielik-PL-11B-v3.0-Instruct against Bielik-11B-v3.0-Instruct, and the two arms of Qwen2.5-1.5B after 500,170,752 Qwen tokens of continued pretraining on the same document sequence; words matched on frequency decile and character length, with within-type variation across contexts; identity specificity required: patching a different word of the same token length must change the emission to that word. Polish texts are the comparison on which the mediated share is predicted near zero, without a fixed bound.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Mediated share

    The average causal mediation effect of a mediator divided by the total effect of a treatment on an outcome, with an item bootstrap interval.

    Mediates at least

    In the pair of pair_1_subject and pair_1_comparator and in the pair of pair_2_subject and pair_2_comparator, the measure of the mediator in the effect of the treatment on the outcome is at least threshold on text_1 and on text_2.

    Word reconstruction depth

    The shallowest layer at which a model's hidden state at the final token of a word in context, patched into a fixed neutral prompt that asks for a word to be repeated, makes the model emit that word's surface form.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    math_clean statements

    Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.

    {
      "wording": "On English Python code and math_clean statements, word reconstruction depth mediates at least 30% of the per-token surprisal penalty of the APT4 model over its original-tokenizer counterpart, in the Bielik 11B v3 pair and in the Qwen2.5-1.5B continued-pretraining pair.",
      "predicate": {
        "type": "concept",
        "key": "mediates_at_least"
      },
      "roles": [
        {
          "role": "measure",
          "definition": "Quantity bounded.",
          "value": {
            "type": "concept",
            "key": "mediated_share"
          }
        },
        {
          "role": "mediator",
          "definition": "Variable carrying the effect.",
          "value": {
            "type": "concept_ref",
            "versionId": "02ed29ad-dcef-43db-a25a-c9914e513112",
            "key": "reconstruction_depth"
          }
        },
        {
          "role": "treatment",
          "definition": "Variable whose effect is decomposed.",
          "value": {
            "type": "text",
            "value": "difference in tokens per word between the pair's tokenizers"
          }
        },
        {
          "role": "outcome",
          "definition": "Variable affected.",
          "value": {
            "type": "text",
            "value": "per-token surprisal penalty of the subject over the comparator"
          }
        },
        {
          "role": "threshold",
          "definition": "Lower bound on the measure.",
          "value": {
            "type": "decimal",
            "value": "0.3",
            "unit": "ratio"
          }
        },
        {
          "role": "pair_1_subject",
          "definition": "APT4 model of the first pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "pair_1_comparator",
          "definition": "Original-tokenizer model of the first pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "pair_2_subject",
          "definition": "APT4 model of the second pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "pair_2_comparator",
          "definition": "Original-tokenizer model of the second pair, after the same continued pretraining.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "text_1",
          "definition": "First text.",
          "value": {
            "type": "concept_ref",
            "versionId": "4acc878a-a4a3-45bd-bb03-2cd8541a7856",
            "key": "english_python_code"
          }
        },
        {
          "role": "text_2",
          "definition": "Second text.",
          "value": {
            "type": "concept_ref",
            "versionId": "1c1e659f-9adf-49a7-8976-649eb293aaa7",
            "key": "math_clean_statements"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →