Immutable accepted plan · retrospective
Stage 1: bits per byte at time zero of random, FVT and FOCUS APT4 transplants of Qwen2.5-1.5B on ten domains
The stage-1 plan as run on 5 July 2026, from the design text dated that day, which was first committed together with the results on 6 July.
Public source
https://github.com/stw2/tokenizer-science-tax @ e54c7e02bb1ea89dbaff6911c87331d705602b1e
Reference checked 2026-09-14 13:12 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- Init damage is concentrated on formal domains for random and FVT initialisation; FOCUS at most FVT at most random holds on nearly every domain; the init-quality spread is concentrated on formal domains; the FVT−FOCUS gap is larger on formal than on prose domains.
- Protocol
- 01 trains fastText auxiliary embeddings on the first 200 MB of Polish training text tokenized with APT4 and freezes the 32,000-row matrix; aux_sanity checks it; 02 builds three transplants that share the base model, APT4 and the handling of overlap, byte and special pieces and differ only in how new-token embeddings are initialised; 03 scores the base model and the three transplants in bits per byte on the first 200 documents of ten domains with 2,048-token windows, storing per-document values; 04 computes init damage, init-quality spread and FVT−FOCUS gap per domain with paired document bootstrap intervals, the four hypothesis bins and gates G1 and G4; 05 writes flat headline values.
- Dataset
- Qwen2.5-1.5B and APT4; the first 200 documents of Polish web text, English web text, Polish reviews, Polish Wikipedia science articles, Polish PES examination questions, English arXiv abstracts, English LaTeX method sections, English Python code and synthetic arithmetic problems, plus GSM8K problems as a contamination reference; the first 200 MB of Polish FineWeb2-HQ training text for the FOCUS auxiliary embeddings.
- Split
- Evaluation holdouts and frozen corpora; no transplant is trained in stage 1. GSM8K problems are excluded from every endpoint because the base model has memorised them.
- Access needs
- The APT4 tokenizer is gated on Hugging Face. The transplant models, the auxiliary corpus, the Polish training text, the English web holdout and four corpora are restricted materials or are rebuilt by script at the pinned revisions.
- Configurations
- Random: normal with the base embedding matrix's global mean and standard deviation. FVT: mean of the constituent Qwen rows. FOCUS: sparsemax-weighted overlap rows by cosine similarity of fastText vectors (skip-gram, 200 dimensions, character n-grams 3 to 6, 3 epochs, 4 workers). Overlap pieces copy their Qwen row in FVT and FOCUS. Scoring: 200 documents, 2,048-token windows, no special tokens, bf16.
- Metric
- Bits per byte per domain; init damage, init-quality spread and FVT−FOCUS gap per domain, the last two with 95% paired document bootstrap intervals of 1000 resamples.
- Seeds
- 20260705 for the random initialisation, the fastText training and the bootstrap.
- Interpretation rule
- H1 passes if the formal-to-prose ratio of mean init damage exceeds 1.15 for both random and FVT initialisation; H2 if FOCUS at most FVT at most random holds on at least 7 of 9 domains; H3 if the formal-to-prose ratio of mean init-quality spread exceeds 1.25; H4 if the mean FVT−FOCUS gap over formal domains exceeds its mean over prose domains. If H1 and H3 both fail, the outcome is a null. Gates: the base model reproduces the full-parameter experiment's time-zero base row within 0.005 bits per byte on 9 domains; the rebuilt FVT transplant's sanity loss is about 10.3; each initialisation rebuilds byte-identically; FVT and FOCUS are equal on overlap rows; vocabulary 32,000 with tied embeddings.
- Resources
- Apple silicon, 128 GB unified memory: CPU for fastText and builds, MPS for scoring, about one hour of scoring; no training.
- Prior work
- arXiv:2604.10799v1 chooses FOCUS citing aggregate results on Bielik 1.5B v3 and reports no per-domain comparison of initialisations.
Selected exact hypotheses and premises
Premise · P24
f5ada7c4-30c8-405f-8741-cd6fcfb58530The aggregate evidence the per-domain comparison complements.
Premise · P21
5a078e71-84ff-4039-9e32-7999a5f679f5The auxiliary embedding space FOCUS needs and the paper does not specify.
Premise · P103
b7e7c5ce-2805-43d7-9f89-ccda038df1b5The independently built FVT transplant of the full-parameter experiment, whose recipe the FVT initialisation follows and whose sanity loss gate G2 compares with.
Premise · P113
faa76c5f-a834-41c9-8e3f-e4422de4c981Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P117
d9ab6218-cf24-498a-90f1-ebd986fb3508Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P118
ee34e809-5b64-4f37-a761-829e94494bf7Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P119
bbd085f6-467a-47b4-934f-83c089d74d9aIts starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P120
82fd08a5-cca0-4136-a0c3-a146e9d7d14eIts starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P122
7ff16749-a07c-49e4-b96e-f230b7c9680cIts starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P114
f831bb88-fc0d-46ff-88cc-fb9e26e29f8eIts starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P115
c5737cda-1d82-4f3a-94fa-e13f3310742dIts starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.
Premise · P116
f81ecf8a-0733-4378-ac05-119e586a18cbIts starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.