Exact proposal revision
How much of the change from an untouched model to the same model after an APT4 transplant and continued pretraining is attributable to the continued pretraining rather than to the tokenizer replacement, when the original-tokenizer model receives identical continued pretraining?
Qwen2.5-1.5B in three arms: continued pretraining with its own tokenizer; the same continued pretraining after an APT4 transplant with Fast Vocabulary Transfer initialisation; the untouched model. 0.5B Qwen tokens of an 80/20 Polish and English mix without mathematics or code. Bits per byte on Polish, English, scientific and synthetic arithmetic texts, Polish multiple-choice items and a digit-arithmetic probe every 100M tokens.
Access and suggested protocol
- Access needs
- The APT4 tokenizer is gated on Hugging Face. Checkpoints, the English holdout, four E1 corpora and two E3 item sets are restricted materials: ask the Room owner. The training text and packed data were not kept.
- Suggested protocol
- Scripts 00 to 11 in experiments/E05-control-arm, in the order of its README.
Selected exact hypotheses and premises
Premise · P12
d5bb75c0-40ed-4d4c-afa4-0fded6fbcb94The intervention the transplant arm reproduces at small scale.
Reason for this revision
Initial proposal.