Exact proposal revision
Does a small amount of post-training on verifiable arithmetic recover more of the APT4 arm's paired digit-probe deficit than math-and-code recovery pretraining does, and does reinforcement learning with a verifiable reward cause less bits-per-byte damage than supervised fine-tuning at the same gain?
A 2 by 3 factorial on the two continued-pretraining arms of Qwen2.5-1.5B after 500M tokens: no post-training, supervised fine-tuning on 5,000 synthetic arithmetic and STEM examples, and GRPO with an exact-match reward, each within 5M tokens. Endpoints: the paired digit probe with a McNemar test; bits per byte on the ten evaluation texts as the forgetting endpoint, with GSM8K problems only as a contamination-decay reference; pass@1 and pass@k on a held-out set written after the models' training data; KL divergence to the policy before post-training; extraction rate; a likelihood-scored arithmetic endpoint. The comparable quantity is the APT4 arm minus the original-tokenizer arm after post-training against the same difference before it.
Access and suggested protocol
- Access needs
- The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Training needs one CUDA GPU with 16 GB.
- Suggested protocol
- Commit a design before training. Format compliance is the feared artifact: any post-training fixes answer formatting, so only the arm-differenced paired quantity is reported as recovery, extraction rate is a separate endpoint, and a likelihood-scored arithmetic endpoint that formatting cannot move is added; pass@k is reported because reinforcement learning can raise pass@1 while lowering pass@k. Training data: scripts/05_gen_probes.py of experiments/E02-digit-policy with a seed disjoint from the probe's; no GSM8K-derived data. Kill criterion: the arm difference after post-training indistinguishable from the difference before it, which makes post-training orthogonal to the transplant's digit deficit.
Selected exact hypotheses and premises
Hypothesis · H51
a6ae6c42-58f8-4e28-866e-b25d3236ec25Recovered share of the digit deficit against the rescue fraction.
Hypothesis · H52
d11ae9a5-9243-4558-b2e9-d15818f5cf90Collateral bits-per-byte damage of the two methods at matched gain.
Premise · P133
683a738d-67cc-4630-b105-e5ad71c2d77aThe paired digit-probe deficit of the APT4 arm at matched continued pretraining.
Premise · P165
09b0327f-40c2-43fb-b6a9-9bbb6fa432daThe digit-probe rescue fraction of math and code in recovery pretraining at the same model size.
Premise · P72
59f84e17-2dbc-4e7a-b481-9fe8fc19f29bThe released 11B pair, which differs by post-training as well as by tokenizer, on GSM8K.
Premise · P29
36da68d9-97cf-4575-ad42-ed56e0e895e9The reported GSM8K difference between the released 11B models.
Premise · P12
d5bb75c0-40ed-4d4c-afa4-0fded6fbcb94The released transplant includes post-training after vocabulary adaptation.
Reason for this revision
Initial proposal.