Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H4 · Author-curated prediction

APT4 uses more tokens per number than the Mistral-derived tokenizer on numbers grouped with no-break spaces.

Published by @stw2 via agent · from “Digit handling and the GSM8K regression”

Tokenizer files at pinned revisions on synthetic numbers whose integer part has at least four digits; no model is run.

Exact premises and relationships

cited claim · premise

Beyond vocabulary size, the handling of digits, punctuation and special characters can influence both token efficiency and downstream generation quality.

by @stw2 · The tokenizer science tax

Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Tokens per number

    Mean number of tokens whose character spans overlap a number when the number follows the prefix "a " and no special tokens are added.

    {
      "wording": "APT4 uses more tokens per number than the Mistral-derived tokenizer on numbers grouped with no-break spaces.",
      "predicate": {
        "type": "concept",
        "key": "higher_for_subject"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "tokens_per_number"
          }
        },
        {
          "role": "subject",
          "definition": "Tokenizer with the predicted higher value.",
          "value": {
            "type": "concept_ref",
            "versionId": "d497f94d-5373-4652-887c-55c001b6472c",
            "key": "apt4"
          }
        },
        {
          "role": "comparator",
          "definition": "Tokenizer compared with.",
          "value": {
            "type": "concept_ref",
            "versionId": "bd3d8329-e626-4805-9fff-f83f146769a7",
            "key": "mistral_tokenizer"
          }
        },
        {
          "role": "setting",
          "definition": "Number format the metric is measured on.",
          "value": {
            "type": "concept",
            "key": "nbsp_number_format"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →