Lower at matched gain
For model_1 and for model_2, at equal gain in matched_on, the measure after the subject method is lower than after the comparator method.
Hypothesis
Sign in with GitHubHypothesis · H52 · Author-curated prediction
Post-training budget at most 5M tokens per method; no GSM8K-derived training data; format compliance is guarded by the arm-differenced quantity, the extraction rate and a likelihood-scored arithmetic endpoint; the matching procedure is fixed in the design.
finding · premise
After 0.5B tokens of the same continued pretraining, digit-probe accuracy is 0.4756 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.5167 for Qwen2.5-1.5B with its own tokenizer, a paired difference of -0.0411 (95% interval -0.0645 to -0.0177, sign-test p 0.0008).The paired digit-probe deficit the post-training targets.
Loading research…
For model_1 and for model_2, at equal gain in matched_on, the measure after the subject method is lower than after the comparator method.
A model's bits per byte after post-training minus its bits per byte before post-training, pooled over the Polish FineWeb2-HQ holdout, Polish Wikipedia science articles, Polish PES examination questions, Polish reviews, the English SlimPajama holdout, English arXiv abstracts, English LaTeX method sections, English Python code and math_clean statements.
Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
{
"wording": "At matched digit-probe accuracy gain, GRPO with an exact-match reward on verifiable arithmetic raises the pooled bits per byte of each continued-pretraining arm of Qwen2.5-1.5B on nine evaluation texts other than GSM8K problems less than supervised fine-tuning does.",
"predicate": {
"type": "concept",
"key": "lower_at_matched_gain"
},
"roles": [
{
"role": "subject",
"definition": "Method predicted to damage less.",
"value": {
"type": "text",
"value": "GRPO with an exact-match reward on verifiable arithmetic"
}
},
{
"role": "comparator",
"definition": "Method compared with.",
"value": {
"type": "text",
"value": "supervised fine-tuning on synthetic arithmetic and STEM examples"
}
},
{
"role": "measure",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "collateral_bpb_increase"
}
},
{
"role": "matched_on",
"definition": "Accuracy whose gain is matched.",
"value": {
"type": "concept_ref",
"versionId": "fbf37827-4b31-402b-8eba-0bb9d16d62c9",
"key": "digit_probe_accuracy"
}
},
{
"role": "model_1",
"definition": "First model post-trained.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "model_2",
"definition": "Second model post-trained.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "continued_pretraining_tokens",
"definition": "Tokens of continued pretraining of both arms at the checkpoints named.",
"value": {
"type": "decimal",
"value": "500170752",
"unit": "tokens"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →