Exact proposal revision
What is the between-run variation of each endpoint reported for the continued-pretraining arms of Qwen2.5-1.5B, and which reported differences are smaller than their minimum detectable effect?
Per arm, three continuations from one shared state at 450M tokens with different data orders for the final 50M tokens, a lower bound on full-run variance, scored on every endpoint of the control-arm experiment; a numerics term from the MPS and CPU parity band of the conformance battery; a per-endpoint uncertainty budget and minimum detectable effect, with the rule that no effect smaller than its band is claimed.
Access and suggested protocol
- Access needs
- The control-arm design keeps optimizer state only in the latest training state, not in the 100M-token milestone checkpoints, and the archive holds neither the optimizer states nor the packed streams: ask the Room owner which training states survive; a shared state at 450M tokens needs a resumed run. The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Six short runs on CUDA GPUs with 16 GB or more.
- Suggested protocol
- Preconditions: a free training GPU after the 1B extension, and the conformance battery's parity band. Commit a design before training. A lower bound on variance can only remove effects, never validate one. Training: scripts/04_train.py of experiments/E05-control-arm with a data-order seed per continuation; scoring with scripts/06_score_checkpoint.sh.
Selected exact hypotheses and premises
Hypothesis · H54
f0980561-b601-47b6-ae28-5305874fd88eThe digit deficit against its minimum detectable effect.
Premise · P128
4a7dff3a-4d95-465e-a43e-bb0df933fc9eA formal-text gap reported from one training run per arm.
Premise · P156
6fe7b629-b8ab-4378-987c-7692f165a3d6Digit-probe accuracy of one arm varies widely between its checkpoints.
Premise · P171
0d8424f7-3e17-4327-b38b-6f375ac39213Digit-probe accuracy of a 1.5B arm changes widely between two checkpoints.
Reason for this revision
Initial proposal.