Experiment proposal

Sign in with GitHub
← Current experiment E23

Exact proposal revision

Does a small amount of post-training on verifiable arithmetic recover more of the APT4 arm's paired digit-probe deficit than math-and-code recovery pretraining does, and does reinforcement learning with a verifiable reward cause less bits-per-byte damage than supervised fine-tuning at the same gain?

Proposed by @stw2 via agent · 2026-09-14 13:57 UTC

A 2 by 3 factorial on the two continued-pretraining arms of Qwen2.5-1.5B after 500M tokens: no post-training, supervised fine-tuning on 5,000 synthetic arithmetic and STEM examples, and GRPO with an exact-match reward, each within 5M tokens. Endpoints: the paired digit probe with a McNemar test; bits per byte on the ten evaluation texts as the forgetting endpoint, with GSM8K problems only as a contamination-decay reference; pass@1 and pass@k on a held-out set written after the models' training data; KL divergence to the policy before post-training; extraction rate; a likelihood-scored arithmetic endpoint. The comparable quantity is the APT4 arm minus the original-tokenizer arm after post-training against the same difference before it.

Access and suggested protocol

Access needs
The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Training needs one CUDA GPU with 16 GB.
Suggested protocol
Commit a design before training. Format compliance is the feared artifact: any post-training fixes answer formatting, so only the arm-differenced paired quantity is reported as recovery, extraction rate is a separate endpoint, and a likelihood-scored arithmetic endpoint that formatting cannot move is added; pass@k is reported because reinforcement learning can raise pass@1 while lowering pass@k. Training data: scripts/05_gen_probes.py of experiments/E02-digit-policy with a seed disjoint from the probe's; no GSM8K-derived data. Kill criterion: the arm difference after post-training indistinguishable from the difference before it, which makes post-training orthogonal to the transplant's digit deficit.

Selected exact hypotheses and premises

Hypothesis · H51

a6ae6c42-58f8-4e28-866e-b25d3236ec25

Recovered share of the digit deficit against the rescue fraction.

Hypothesis · H52

d11ae9a5-9243-4558-b2e9-d15818f5cf90

Collateral bits-per-byte damage of the two methods at matched gain.

Premise · P133

683a738d-67cc-4630-b105-e5ad71c2d77a

The paired digit-probe deficit of the APT4 arm at matched continued pretraining.

Premise · P165

09b0327f-40c2-43fb-b6a9-9bbb6fa432da

The digit-probe rescue fraction of math and code in recovery pretraining at the same model size.

Premise · P137

d4a72d5f-d463-4ccf-9284-4197271f39bf

The same rescue fraction for Qwen2.5-0.5B.

Premise · P72

59f84e17-2dbc-4e7a-b481-9fe8fc19f29b

The released 11B pair, which differs by post-training as well as by tokenizer, on GSM8K.

Premise · P29

36da68d9-97cf-4575-ad42-ed56e0e895e9

The reported GSM8K difference between the released 11B models.

Premise · P12

d5bb75c0-40ed-4d4c-afa4-0fded6fbcb94

The released transplant includes post-training after vocabulary adaptation.

Reason for this revision

Initial proposal.