Experiment · E23
Does a small amount of post-training on verifiable arithmetic recover more of the APT4 arm's paired digit-probe deficit than math-and-code recovery pretraining does, and does reinforcement learning with a verifiable reward cause less bits-per-byte damage than supervised fine-tuning at the same gain?
A 2 by 3 factorial on the two continued-pretraining arms of Qwen2.5-1.5B after 500M tokens: no post-training, supervised fine-tuning on 5,000 synthetic arithmetic and STEM examples, and GRPO with an exact-match reward, each within 5M tokens. Endpoints: the paired digit probe with a McNemar test; bits per byte on the ten evaluation texts as the forgetting endpoint, with GSM8K problems only as a contamination-decay reference; pass@1 and pass@k on a held-out set written after the models' training data; KL divergence to the policy before post-training; extraction rate; a likelihood-scored arithmetic endpoint. The comparable quantity is the APT4 arm minus the original-tokenizer arm after post-training against the same difference before it.
Prerequisites and protocol
- Access needs
- The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Training needs one CUDA GPU with 16 GB.
- Suggested protocol
- Commit a design before training. Format compliance is the feared artifact: any post-training fixes answer formatting, so only the arm-differenced paired quantity is reported as recovery, extraction rate is a separate endpoint, and a likelihood-scored arithmetic endpoint that formatting cannot move is added; pass@k is reported because reinforcement learning can raise pass@1 while lowering pass@k. Training data: scripts/05_gen_probes.py of experiments/E02-digit-policy with a seed disjoint from the probe's; no GSM8K-derived data. Kill criterion: the arm difference after post-training indistinguishable from the difference before it, which makes post-training orthogonal to the transplant's digit deficit.
Accepted plan
No accepted plan. Available work need not have a complete protocol or source commit.
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
No attempt registered. Work status and findings are independent of attempts.
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →