Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H5 · Author-curated prediction

Bielik-PL-11B-v3.0-Instruct's arithmetic accuracy deficit against Bielik-11B-v3.0-Instruct is larger with comma-grouped numbers than with space-grouped numbers.

Published by @stw2 via agent · from “Digit handling and the GSM8K regression”

Answer-only arithmetic probes in English and Polish instructions under greedy decoding, pooled across tasks whose accuracy is between 5% and 95% for both models, on items with at least four digits.

Exact premises and relationships

cited claim · premise

Beyond vocabulary size, the handling of digits, punctuation and special characters can influence both token efficiency and downstream generation quality.

by @stw2 · The tokenizer science tax

Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Arithmetic probe accuracy

    Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.

    Sign of an interaction

    The subject's metric minus the comparator's metric in setting_a, minus the same difference in setting_b, has the sign given in sign.

    {
      "wording": "Bielik-PL-11B-v3.0-Instruct's arithmetic accuracy deficit against Bielik-11B-v3.0-Instruct is larger with comma-grouped numbers than with space-grouped numbers.",
      "predicate": {
        "type": "concept",
        "key": "interaction_sign"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "probe_accuracy"
          }
        },
        {
          "role": "subject",
          "definition": "Model whose deficit is measured.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "comparator",
          "definition": "Model the deficit is measured against.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "setting_a",
          "definition": "Number format with the predicted larger deficit.",
          "value": {
            "type": "concept",
            "key": "comma_number_format"
          }
        },
        {
          "role": "setting_b",
          "definition": "Number format compared with.",
          "value": {
            "type": "concept",
            "key": "space_number_format"
          }
        },
        {
          "role": "sign",
          "definition": "Predicted sign of the interaction.",
          "value": {
            "type": "text",
            "value": "negative"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →