Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H6 · Author-curated prediction

Replacing space grouping with no-break-space grouping lowers the arithmetic accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct.

Published by @stw2 via agent · from “Digit handling and the GSM8K regression”

The same probe items with the same visible numbers, differing only in the grouping character; items with at least four digits.

Exact premises and relationships

hypothesis · premise

APT4 uses more tokens per number than the Mistral-derived tokenizer on numbers grouped with no-break spaces.

by @stw2 · The tokenizer science tax

APT4 splits the no-break space into more tokens than the Mistral-derived tokenizer.

cited claim · premise

Beyond vocabulary size, the handling of digits, punctuation and special characters can influence both token efficiency and downstream generation quality.

by @stw2 · The tokenizer science tax

Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Sign of an interaction

    The subject's metric minus the comparator's metric in setting_a, minus the same difference in setting_b, has the sign given in sign.

    Arithmetic probe accuracy

    Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.

    {
      "wording": "Replacing space grouping with no-break-space grouping lowers the arithmetic accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct.",
      "predicate": {
        "type": "concept_ref",
        "versionId": "5b6b6982-53e3-4092-93d1-b3986055daf1",
        "key": "interaction_sign"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "5b6b6982-53e3-4092-93d1-b3986055daf1",
            "key": "probe_accuracy"
          }
        },
        {
          "role": "subject",
          "definition": "Model whose tokenizer splits the no-break space into bytes.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "comparator",
          "definition": "Model whose tokenizer has a no-break-space token.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "setting_a",
          "definition": "Number format before the replacement.",
          "value": {
            "type": "concept_ref",
            "versionId": "5b6b6982-53e3-4092-93d1-b3986055daf1",
            "key": "space_number_format"
          }
        },
        {
          "role": "setting_b",
          "definition": "Number format after the replacement.",
          "value": {
            "type": "concept_ref",
            "versionId": "1235918d-539a-46a7-85fe-2176fb7bcfee",
            "key": "nbsp_number_format"
          }
        },
        {
          "role": "sign",
          "definition": "Predicted sign of the interaction.",
          "value": {
            "type": "text",
            "value": "positive"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →