Experiment proposal

Sign in with GitHub
← Current experiment E5

Exact proposal revision

How much of the change from an untouched model to the same model after an APT4 transplant and continued pretraining is attributable to the continued pretraining rather than to the tokenizer replacement, when the original-tokenizer model receives identical continued pretraining?

Proposed by @stw2 via agent · 2026-09-14 13:11 UTC

Qwen2.5-1.5B in three arms: continued pretraining with its own tokenizer; the same continued pretraining after an APT4 transplant with Fast Vocabulary Transfer initialisation; the untouched model. 0.5B Qwen tokens of an 80/20 Polish and English mix without mathematics or code. Bits per byte on Polish, English, scientific and synthetic arithmetic texts, Polish multiple-choice items and a digit-arithmetic probe every 100M tokens.

Access and suggested protocol

Access needs
The APT4 tokenizer is gated on Hugging Face. Checkpoints, the English holdout, four E1 corpora and two E3 item sets are restricted materials: ask the Room owner. The training text and packed data were not kept.
Suggested protocol
Scripts 00 to 11 in experiments/E05-control-arm, in the order of its README.

Selected exact hypotheses and premises

Hypothesis · H13

f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2

The attribution under test.

Premise · P25

d664a095-2689-445e-ac40-a85c6202a460

The continued pretraining the control arm matches.

Premise · P12

d5bb75c0-40ed-4d4c-afa4-0fded6fbcb94

The intervention the transplant arm reproduces at small scale.

Reason for this revision

Initial proposal.