Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H26 · Author-curated prediction

With the no-break space, narrow no-break space and minus sign added as fixed pieces, the Polish-only 32k tokenizer encodes each of them as one token between digits at a pooled relative token cost below 0.005 on three Polish corpora.

Published by @stw2 via agent · from “Fertility and vocabulary allocation”

Digit and separator tokenization on synthetic number renderings and tokenizer fertility on three Polish corpora; no model is trained or run.

Exact premises and relationships

finding · premise

On numbers of at least four digits grouped with no-break spaces, APT4 uses 10.4214 tokens per number and the Mistral-derived tokenizer 9.0303.

by @stw2 · The tokenizer science tax

APT4 needs more tokens than the Mistral-derived tokenizer on numbers grouped with no-break spaces.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    One token at a cost below a threshold

    Each of the characters encodes as one token between digits with the subject tokenizer, and the subject's metric against the comparator on the evaluation texts pooled is below threshold.

    Polish-only 32k tokenizer with separator pieces

    SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer on the same Polish text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

    Polish-only 32k tokenizer

    SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

    Relative token cost

    Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.

    {
      "wording": "With the no-break space, narrow no-break space and minus sign added as fixed pieces, the Polish-only 32k tokenizer encodes each of them as one token between digits at a pooled relative token cost below 0.005 on three Polish corpora.",
      "predicate": {
        "type": "concept",
        "key": "one_token_at_cost_below"
      },
      "roles": [
        {
          "role": "subject",
          "definition": "Tokenizer with the fixed pieces.",
          "value": {
            "type": "concept",
            "key": "polish_only_separator_32k_tokenizer"
          }
        },
        {
          "role": "comparator",
          "definition": "Tokenizer without them, trained on the same text.",
          "value": {
            "type": "concept_ref",
            "versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
            "key": "polish_only_32k_tokenizer"
          }
        },
        {
          "role": "characters",
          "definition": "Characters given fixed pieces.",
          "value": {
            "type": "text",
            "value": "U+00A0 no-break space; U+202F narrow no-break space; U+2212 minus sign"
          }
        },
        {
          "role": "metric",
          "definition": "Cost compared with the threshold.",
          "value": {
            "type": "concept_ref",
            "versionId": "2be4b5ff-86ee-4e87-b3a4-77932559f05f",
            "key": "relative_token_cost"
          }
        },
        {
          "role": "threshold",
          "definition": "Upper bound on the pooled cost.",
          "value": {
            "type": "decimal",
            "value": "0.005",
            "unit": "ratio"
          }
        },
        {
          "role": "evaluation_text",
          "definition": "Corpora the cost is measured on.",
          "value": {
            "type": "text",
            "value": "Polish PES examination questions; Polish Wikipedia science articles; Polish reviews"
          }
        },
        {
          "role": "language",
          "definition": "Language of the corpora.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "polish"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →