Immutable accepted plan · retrospective
Matched-budget continued pretraining with probe battery v2 and the arm-B tokenizer guard
Adds the note of 5 July 2026 (sci_gsm8k is contaminated and excluded from inference), Amendment 1 of 6 July 2026 (a frozen generative digit probe as the primary mathematics endpoint and an arm-B tokenizer-loading fix) and the loader rewrite of 7 July 2026, reconstructed from file times and logs. Written after arm A's 100M to 400M scores and before any arm-B measurement; first committed on 19 July 2026.
Public source
https://github.com/stw2/tokenizer-science-tax @ b4c55e6e7d26917e239ccbe2f11886c4d4a733a2
Reference checked 2026-09-14 13:10 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- A substantial share of the change is attributable to continued pretraining; no threshold was fixed.
- Protocol
- 01 streams FineWeb2-HQ and SlimPajama-6B with a fixed shuffle, holds out 2,000 documents of each and counts tokens with both tokenizers; 03 builds the arm-B transplant with Fast Vocabulary Transfer; 02 packs the shared document sequence for each arm; 04 runs a smoke training, then trains arm A and arm B to 0.5B Qwen tokens with milestones every 100M tokens; 05 scores bits per byte on ten texts (first 200 documents each) and 07 scores multiple-choice accuracy for the base, the transplant at time zero and every milestone. The port adds 00, which fetches the APT4 tokenizer files, and 06, which drives 05, 07 and 08 per model. Amendment 1: 08 freezes a digit-probe spec of 300 E2 arithmetic items in three prompt formats and scores every model with greedy 4-shot completions graded by E2's grader; arm-B tokenizers load through a whitespace canary and the frozen APT4 reference.
- Dataset
- Training: about 80% FineWeb2-HQ pol_Latn and 20% SlimPajama-6B English by Qwen tokens, one deterministic document sequence for both arms. Evaluation: the first 200 of 2,000 held-out documents of each training source, the first 200 documents of seven E1 corpora, the first 200 math_clean statements, and the first 300 LLMzSzŁ STEM and 250 PES multiple-choice items from E3.
- Split
- The first 2,000 documents of each seed-42 shuffled source stream are held out and never trained on; training reads the remaining documents in one interleaved order.
- Access needs
- The APT4 tokenizer is gated on Hugging Face. Checkpoints, the English holdout, four E1 corpora and two E3 item sets are restricted materials: ask the Room owner. The training text and packed data were not kept.
- Configurations
- Arm A: Qwen2.5-1.5B with its own tokenizer. Arm B: APT4 transplanted with Fast Vocabulary Transfer (piece embeddings as means of the constituent Qwen rows; <s> and </s> from <|endoftext|>; <unk> and 256 byte pieces from the mean embedding; tied embeddings; eos 2). Arm C: untouched. Arms A and B: bf16 full fine-tune, sequence 2048, micro-batch 8, gradient accumulation 32 (524,288 tokens per step), 8-bit AdamW with betas 0.9 and 0.95, weight decay 0.1, clip 1.0, 50 warmup steps then constant learning rate 1e-4, milestones every 1e8 tokens.
- Metric
- Bits per UTF-8 byte of each text per milestone; gap = arm B minus arm A; continued-pretraining effect = arm A minus arm C. Pooled digit-probe accuracy is the primary downstream mathematics endpoint, multiple-choice accuracy the knowledge endpoint and math_clean bits per byte the compression endpoint.
- Seeds
- Stream shuffle seed 42 with a 10,000-document buffer; packing and training use no random numbers; one run per arm.
- Interpretation rule
- None fixed; substantial is not defined. The campaign plan: arms A and B about equal on scientific bits per byte would mean the transplant cost is mostly a continued-pretraining effect. sci_gsm8k is a contamination-decay reference only and is excluded from inference.
- Resources
- Training on one CUDA GPU, 16 GB, arms serial, about 4 days per arm at 0.5B tokens by the design's estimate; scoring on a separate machine, Apple silicon, 128 GB unified memory.
- Prior work
- arXiv:2604.10799v1 compares the transplanted models with the untouched originals only; no arm continues the original tokenizer on the same tokens.
Selected exact hypotheses and premises
Premise · P25
d664a095-2689-445e-ac40-a85c6202a460The continued pretraining the control arm matches at small scale.
Premise · P22
01087ca0-917c-46f8-a25e-ebc786018d72The embedding initialisation of the transplant arm.