Immutable accepted plan · retrospective
Matched-budget continued pretraining of Qwen2.5-1.5B with its own tokenizer and after an APT4 transplant, against the untouched model
The design as run from 4 July 2026. It was first committed together with both amendments and every result on 19 July 2026, ten days after the last measurement, so the plan is retrospective.
Public source
https://github.com/stw2/tokenizer-science-tax @ 4d64dedcbe57761c48ea37d5ce80cb6c86198236
Reference checked 2026-09-14 13:10 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- A substantial share of the change is attributable to continued pretraining; no threshold was fixed.
- Protocol
- 01 streams FineWeb2-HQ and SlimPajama-6B with a fixed shuffle, holds out 2,000 documents of each and counts tokens with both tokenizers; 03 builds the arm-B transplant with Fast Vocabulary Transfer; 02 packs the shared document sequence for each arm; 04 runs a smoke training, then trains arm A and arm B to 0.5B Qwen tokens with milestones every 100M tokens; 05 scores bits per byte on ten texts (first 200 documents each) and 07 scores multiple-choice accuracy for the base, the transplant at time zero and every milestone. The port adds 00, which fetches the APT4 tokenizer files, and 06, which drives 05, 07 and 08 per model.
- Dataset
- Training: about 80% FineWeb2-HQ pol_Latn and 20% SlimPajama-6B English by Qwen tokens, one deterministic document sequence for both arms. Evaluation: the first 200 of 2,000 held-out documents of each training source, the first 200 documents of seven E1 corpora, the first 200 math_clean statements, and the first 300 LLMzSzŁ STEM and 250 PES multiple-choice items from E3.
- Split
- The first 2,000 documents of each seed-42 shuffled source stream are held out and never trained on; training reads the remaining documents in one interleaved order.
- Access needs
- The APT4 tokenizer is gated on Hugging Face. Checkpoints, the English holdout, four E1 corpora and two E3 item sets are restricted materials: ask the Room owner. The training text and packed data were not kept.
- Configurations
- Arm A: Qwen2.5-1.5B with its own tokenizer. Arm B: APT4 transplanted with Fast Vocabulary Transfer (piece embeddings as means of the constituent Qwen rows; <s> and </s> from <|endoftext|>; <unk> and 256 byte pieces from the mean embedding; tied embeddings; eos 2). Arm C: untouched. Arms A and B: bf16 full fine-tune, sequence 2048, micro-batch 8, gradient accumulation 32 (524,288 tokens per step), 8-bit AdamW with betas 0.9 and 0.95, weight decay 0.1, clip 1.0, 50 warmup steps then constant learning rate 1e-4, milestones every 1e8 tokens.
- Metric
- Bits per UTF-8 byte of each text per milestone; gap = arm B minus arm A; continued-pretraining effect = arm A minus arm C; multiple-choice accuracy as a downstream probe.
- Seeds
- Stream shuffle seed 42 with a 10,000-document buffer; packing and training use no random numbers; one run per arm.
- Interpretation rule
- None fixed; substantial is not defined. The campaign plan: arms A and B about equal on scientific bits per byte would mean the transplant cost is mostly a continued-pretraining effect.
- Resources
- Training on one CUDA GPU, 16 GB, arms serial, about 4 days per arm at 0.5B tokens by the design's estimate; scoring on a separate machine, Apple silicon, 128 GB unified memory.
- Prior work
- arXiv:2604.10799v1 compares the transplanted models with the untouched originals only; no arm continues the original tokenizer on the same tokens.
Selected exact hypotheses and premises
Premise · P25
d664a095-2689-445e-ac40-a85c6202a460The continued pretraining the control arm matches at small scale.
Premise · P22
01087ca0-917c-46f8-a25e-ebc786018d72The embedding initialisation of the transplant arm.