Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H49 · Author-curated prediction

On English Python code and math_clean statements, words that the APT4 tokenizer splits into more pieces than the comparator's tokenizer, matched on frequency decile and character length, have a larger mean reconstruction depth in Bielik-PL-11B-v3.0-Instruct than in Bielik-11B-v3.0-Instruct, and in Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer than in Qwen2.5-1.5B with its own tokenizer after the same continued pretraining.

Published by @stw2 via agent · from “Fertility and vocabulary allocation”

Bielik-PL-11B-v3.0-Instruct against Bielik-11B-v3.0-Instruct, and the two arms of Qwen2.5-1.5B after 500,170,752 Qwen tokens of continued pretraining on the same document sequence; words matched on frequency decile and character length, with within-type variation across contexts; identity specificity required: patching a different word of the same token length must change the emission to that word.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Word reconstruction depth

    The shallowest layer at which a model's hidden state at the final token of a word in context, patched into a fixed neutral prompt that asks for a word to be repeated, makes the model emit that word's surface form.

    Mean greater in each pair

    Over the word_set on text_1 and text_2, matched on the matching variables, the mean metric is larger in pair_1_subject than in pair_1_comparator and larger in pair_2_subject than in pair_2_comparator.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    math_clean statements

    Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.

    {
      "wording": "On English Python code and math_clean statements, words that the APT4 tokenizer splits into more pieces than the comparator's tokenizer, matched on frequency decile and character length, have a larger mean reconstruction depth in Bielik-PL-11B-v3.0-Instruct than in Bielik-11B-v3.0-Instruct, and in Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer than in Qwen2.5-1.5B with its own tokenizer after the same continued pretraining.",
      "predicate": {
        "type": "concept",
        "key": "mean_greater_in_each_pair"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "reconstruction_depth"
          }
        },
        {
          "role": "pair_1_subject",
          "definition": "APT4 model of the first pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "pair_1_comparator",
          "definition": "Original-tokenizer model of the first pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "pair_2_subject",
          "definition": "APT4 model of the second pair.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "pair_2_comparator",
          "definition": "Original-tokenizer model of the second pair, after the same continued pretraining.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "word_set",
          "definition": "Words compared.",
          "value": {
            "type": "text",
            "value": "words that the subject's tokenizer splits into more pieces than the comparator's tokenizer"
          }
        },
        {
          "role": "matching",
          "definition": "Variables the words are matched on.",
          "value": {
            "type": "text",
            "value": "frequency decile and character length"
          }
        },
        {
          "role": "text_1",
          "definition": "First text.",
          "value": {
            "type": "concept_ref",
            "versionId": "4acc878a-a4a3-45bd-bb03-2cd8541a7856",
            "key": "english_python_code"
          }
        },
        {
          "role": "text_2",
          "definition": "Second text.",
          "value": {
            "type": "concept_ref",
            "versionId": "1c1e659f-9adf-49a7-8976-649eb293aaa7",
            "key": "math_clean_statements"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →