Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H25 · Author-curated prediction

The science-slice 32k tokenizer's relative token cost against the Polish-only 32k tokenizer is below 0.05 on three Polish corpora pooled and below 0.07 on each of them.

Published by @stw2 via agent · from “Fertility and vocabulary allocation”

Tokenizer fertility on three Polish corpora; no model is trained or run.

Exact premises and relationships

cited claim · premise

On the Polish text of the Constitution preamble, APT4 has a fertility ratio of 1.62 tokens per word and the Mistral-derived tokenizer 3.22.

by @stw2 · The tokenizer science tax

The Polish fertility gain of a Polish-optimised 32k tokenizer that a science slice could erode.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Relative token cost

    Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.

    Cost below thresholds

    The subject's metric against the comparator is below pooled_threshold on the evaluation texts pooled and below per_text_threshold on each of them.

    Polish-only 32k tokenizer

    SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

    Science-slice 32k tokenizer

    SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

    {
      "wording": "The science-slice 32k tokenizer's relative token cost against the Polish-only 32k tokenizer is below 0.05 on three Polish corpora pooled and below 0.07 on each of them.",
      "predicate": {
        "type": "concept",
        "key": "cost_below_thresholds"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared with the thresholds.",
          "value": {
            "type": "concept",
            "key": "relative_token_cost"
          }
        },
        {
          "role": "subject",
          "definition": "Tokenizer whose cost is predicted.",
          "value": {
            "type": "concept_ref",
            "versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
            "key": "science_slice_32k_tokenizer"
          }
        },
        {
          "role": "comparator",
          "definition": "Tokenizer the cost is measured against.",
          "value": {
            "type": "concept_ref",
            "versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
            "key": "polish_only_32k_tokenizer"
          }
        },
        {
          "role": "evaluation_text",
          "definition": "Corpora the cost is measured on.",
          "value": {
            "type": "text",
            "value": "Polish PES examination questions; Polish Wikipedia science articles; Polish reviews"
          }
        },
        {
          "role": "pooled_threshold",
          "definition": "Upper bound on the pooled cost.",
          "value": {
            "type": "decimal",
            "value": "0.05",
            "unit": "ratio"
          }
        },
        {
          "role": "per_text_threshold",
          "definition": "Upper bound on the cost on each corpus.",
          "value": {
            "type": "decimal",
            "value": "0.07",
            "unit": "ratio"
          }
        },
        {
          "role": "language",
          "definition": "Language of the corpora.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "polish"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →