Exact proposal revision
Is the bits-per-byte residual of the APT4 arm over the original-tokenizer arm on English and formal texts permanent, or does it keep closing when both continued-pretraining arms of Qwen2.5-1.5B continue to 1B tokens and are annealed?
Both continued-pretraining arms of Qwen2.5-1.5B resumed from their constant-learning-rate state at 500,170,752 Qwen tokens to 1B tokens on the same document sequence, which stays under one epoch, followed by the WSD decay anneal of both final checkpoints; bits per byte on the ten evaluation texts, the digit probe and the multiple-choice probes at every 100M-token milestone and after the anneal. Training and scoring.
Access and suggested protocol
- Access needs
- An exact resume needs each arm's optimizer and data-sampler state and the packed training streams; none is among the archived materials, and the digests of the training text and packed streams were not recorded: ask the Room owner whether they survive, otherwise rebuild the streams with scripts 01 and 02 and record that the resume is not exact. The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Training needs one CUDA GPU with 16 GB, serially, about 1.5 days per arm.
- Suggested protocol
- Amendment 2 of the control-arm design, adopted on 8 July 2026 and never launched, as its own experiment. Commit a design with a decision rule before resuming. Training: scripts/04_train.py of experiments/E05-control-arm with --target-tokens 1e9 for the original-tokenizer arm and then the APT4 arm, each from its run directory, then --anneal on both final checkpoints before any headline value; the APT4 pack caps that arm near 1.0B tokens. Scoring: scripts/06_score_checkpoint.sh at each milestone; byte-anchored windows are added alongside the canonical windows once the per-position scoring rig exists, and checkpoints stay re-scorable, so the extension does not wait for it.
Selected exact hypotheses and premises
Premise · P125
f1355715-5d7f-4a04-a4bf-249f33aa288eThe English gap after each 100M tokens, still falling at 500M.
Premise · P126
b11964d8-81ba-4bf6-9364-302e0e5a677cThe arXiv-abstract gap after each 100M tokens, still falling at 500M.
Premise · P127
ebd85d87-d358-44c8-aa85-5e8246d99b2fThe LaTeX gap after each 100M tokens, still falling at 500M.
Premise · P128
4a7dff3a-4d95-465e-a43e-bb0df933fc9eThe Python gap after each 100M tokens, still falling at 500M.
Premise · P129
e2da0e41-7694-4b22-a1ea-14d13dc51f38The math_clean gap after each 100M tokens, level between 400M and 500M.
Premise · P124
131e534a-1b09-40c1-95fc-a4d5d63faae0The Polish gap after each 100M tokens, near closed at 500M.
Reason for this revision
Initial proposal.