Exceeds a reported value
For each of the methods, applied to subject and comparator within post_training_budget, the measure on the probe exceeds the value reported in reference_value.
Hypothesis
Sign in with GitHubHypothesis · H51 · Author-curated prediction
Both arms after 500,170,752 Qwen tokens of continued pretraining; no GSM8K-derived training data; format compliance is guarded by the arm-differenced quantity, the extraction rate and a likelihood-scored arithmetic endpoint.
finding · premise
After 0.5B tokens of the same continued pretraining, digit-probe accuracy is 0.4756 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.5167 for Qwen2.5-1.5B with its own tokenizer, a paired difference of -0.0411 (95% interval -0.0645 to -0.0177, sign-test p 0.0008).The paired digit-probe deficit of the APT4 arm.
finding · premise
After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-1.5B, the rescue fraction of a 30% math and code share on digit-probe accuracy is 0.4706 (95% interval 0.3571 to 0.5871); the gap to the base model is 0.1889 without math and code and 0.1000 with the 30% share.The digit-probe rescue fraction of math and code in recovery pretraining at the same model size.
Loading research…
For each of the methods, applied to subject and comparator within post_training_budget, the measure on the probe exceeds the value reported in reference_value.
One minus the ratio of the subject's accuracy minus the comparator's accuracy after the same post-training of both to the subject's accuracy minus the comparator's accuracy before post-training, on the same items.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
{
"wording": "Post-training of at most 5M tokens on verifiable arithmetic, by supervised fine-tuning or by GRPO with an exact-match reward, applied to both continued-pretraining arms of Qwen2.5-1.5B recovers a larger share of the APT4 arm's paired digit-probe accuracy deficit than the rescue fraction of a 30% math and code share in 1B tokens of recovery pretraining of an APT4 transplant of Qwen2.5-1.5B.",
"predicate": {
"type": "concept",
"key": "exceeds_reported_value"
},
"roles": [
{
"role": "subject",
"definition": "Arm with the deficit.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "comparator",
"definition": "Arm the deficit is measured against.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "continued_pretraining_tokens",
"definition": "Tokens of continued pretraining of both arms at the checkpoints named.",
"value": {
"type": "decimal",
"value": "500170752",
"unit": "tokens"
}
},
{
"role": "methods",
"definition": "Post-training methods, each applied to both arms.",
"value": {
"type": "text",
"value": "supervised fine-tuning on synthetic arithmetic and STEM examples; GRPO with an exact-match reward"
}
},
{
"role": "post_training_budget",
"definition": "Upper bound on post-training tokens per method.",
"value": {
"type": "decimal",
"value": "5000000",
"unit": "tokens"
}
},
{
"role": "probe",
"definition": "Accuracy the deficit is measured on.",
"value": {
"type": "concept_ref",
"versionId": "fbf37827-4b31-402b-8eba-0bb9d16d62c9",
"key": "digit_probe_accuracy"
}
},
{
"role": "measure",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "post_training_recovered_share"
}
},
{
"role": "reference_value",
"definition": "Record reporting the value exceeded.",
"value": {
"type": "record",
"versionId": "09b0327f-40c2-43fb-b6a9-9bbb6fa432da"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →