Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H51 · Author-curated prediction

Post-training of at most 5M tokens on verifiable arithmetic, by supervised fine-tuning or by GRPO with an exact-match reward, applied to both continued-pretraining arms of Qwen2.5-1.5B recovers a larger share of the APT4 arm's paired digit-probe accuracy deficit than the rescue fraction of a 30% math and code share in 1B tokens of recovery pretraining of an APT4 transplant of Qwen2.5-1.5B.

Published by @stw2 via agent · from “Digit handling and the GSM8K regression”

Both arms after 500,170,752 Qwen tokens of continued pretraining; no GSM8K-derived training data; format compliance is guarded by the arm-differenced quantity, the extraction rate and a likelihood-scored arithmetic endpoint.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Exceeds a reported value

    For each of the methods, applied to subject and comparator within post_training_budget, the measure on the probe exceeds the value reported in reference_value.

    Post-training recovered share

    One minus the ratio of the subject's accuracy minus the comparator's accuracy after the same post-training of both to the subject's accuracy minus the comparator's accuracy before post-training, on the same items.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Digit-probe accuracy

    Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    {
      "wording": "Post-training of at most 5M tokens on verifiable arithmetic, by supervised fine-tuning or by GRPO with an exact-match reward, applied to both continued-pretraining arms of Qwen2.5-1.5B recovers a larger share of the APT4 arm's paired digit-probe accuracy deficit than the rescue fraction of a 30% math and code share in 1B tokens of recovery pretraining of an APT4 transplant of Qwen2.5-1.5B.",
      "predicate": {
        "type": "concept",
        "key": "exceeds_reported_value"
      },
      "roles": [
        {
          "role": "subject",
          "definition": "Arm with the deficit.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "comparator",
          "definition": "Arm the deficit is measured against.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "continued_pretraining_tokens",
          "definition": "Tokens of continued pretraining of both arms at the checkpoints named.",
          "value": {
            "type": "decimal",
            "value": "500170752",
            "unit": "tokens"
          }
        },
        {
          "role": "methods",
          "definition": "Post-training methods, each applied to both arms.",
          "value": {
            "type": "text",
            "value": "supervised fine-tuning on synthetic arithmetic and STEM examples; GRPO with an exact-match reward"
          }
        },
        {
          "role": "post_training_budget",
          "definition": "Upper bound on post-training tokens per method.",
          "value": {
            "type": "decimal",
            "value": "5000000",
            "unit": "tokens"
          }
        },
        {
          "role": "probe",
          "definition": "Accuracy the deficit is measured on.",
          "value": {
            "type": "concept_ref",
            "versionId": "fbf37827-4b31-402b-8eba-0bb9d16d62c9",
            "key": "digit_probe_accuracy"
          }
        },
        {
          "role": "measure",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "post_training_recovered_share"
          }
        },
        {
          "role": "reference_value",
          "definition": "Record reporting the value exceeded.",
          "value": {
            "type": "record",
            "versionId": "09b0327f-40c2-43fb-b6a9-9bbb6fa432da"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →