Accepted plan

Sign in with GitHub
← Experiment E7

Immutable accepted plan · retrospective

Stage 1: bits per byte at time zero of random, FVT and FOCUS APT4 transplants of Qwen2.5-1.5B on ten domains

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

The stage-1 plan as run on 5 July 2026, from the design text dated that day, which was first committed together with the results on 6 July.

Public source

Plan

Prediction
Init damage is concentrated on formal domains for random and FVT initialisation; FOCUS at most FVT at most random holds on nearly every domain; the init-quality spread is concentrated on formal domains; the FVT−FOCUS gap is larger on formal than on prose domains.
Protocol
01 trains fastText auxiliary embeddings on the first 200 MB of Polish training text tokenized with APT4 and freezes the 32,000-row matrix; aux_sanity checks it; 02 builds three transplants that share the base model, APT4 and the handling of overlap, byte and special pieces and differ only in how new-token embeddings are initialised; 03 scores the base model and the three transplants in bits per byte on the first 200 documents of ten domains with 2,048-token windows, storing per-document values; 04 computes init damage, init-quality spread and FVT−FOCUS gap per domain with paired document bootstrap intervals, the four hypothesis bins and gates G1 and G4; 05 writes flat headline values.
Dataset
Qwen2.5-1.5B and APT4; the first 200 documents of Polish web text, English web text, Polish reviews, Polish Wikipedia science articles, Polish PES examination questions, English arXiv abstracts, English LaTeX method sections, English Python code and synthetic arithmetic problems, plus GSM8K problems as a contamination reference; the first 200 MB of Polish FineWeb2-HQ training text for the FOCUS auxiliary embeddings.
Split
Evaluation holdouts and frozen corpora; no transplant is trained in stage 1. GSM8K problems are excluded from every endpoint because the base model has memorised them.
Access needs
The APT4 tokenizer is gated on Hugging Face. The transplant models, the auxiliary corpus, the Polish training text, the English web holdout and four corpora are restricted materials or are rebuilt by script at the pinned revisions.
Configurations
Random: normal with the base embedding matrix's global mean and standard deviation. FVT: mean of the constituent Qwen rows. FOCUS: sparsemax-weighted overlap rows by cosine similarity of fastText vectors (skip-gram, 200 dimensions, character n-grams 3 to 6, 3 epochs, 4 workers). Overlap pieces copy their Qwen row in FVT and FOCUS. Scoring: 200 documents, 2,048-token windows, no special tokens, bf16.
Metric
Bits per byte per domain; init damage, init-quality spread and FVT−FOCUS gap per domain, the last two with 95% paired document bootstrap intervals of 1000 resamples.
Seeds
20260705 for the random initialisation, the fastText training and the bootstrap.
Interpretation rule
H1 passes if the formal-to-prose ratio of mean init damage exceeds 1.15 for both random and FVT initialisation; H2 if FOCUS at most FVT at most random holds on at least 7 of 9 domains; H3 if the formal-to-prose ratio of mean init-quality spread exceeds 1.25; H4 if the mean FVT−FOCUS gap over formal domains exceeds its mean over prose domains. If H1 and H3 both fail, the outcome is a null. Gates: the base model reproduces the full-parameter experiment's time-zero base row within 0.005 bits per byte on 9 domains; the rebuilt FVT transplant's sanity loss is about 10.3; each initialisation rebuilds byte-identically; FVT and FOCUS are equal on overlap rows; vocabulary 32,000 with tied embeddings.
Resources
Apple silicon, 128 GB unified memory: CPU for fastText and builds, MPS for scoring, about one hour of scoring; no training.
Prior work
arXiv:2604.10799v1 chooses FOCUS citing aggregate results on Bielik 1.5B v3 and reports no per-domain comparison of initialisations.

Selected exact hypotheses and premises

Hypothesis · H18

c8bc4afe-d6ab-4475-b668-4f01d4d149b2

Formal-damage prediction.

Hypothesis · H19

2174054b-06f0-40a7-b2f7-96cecb0b7c75

Ranking prediction.

Hypothesis · H20

3d170e90-be69-449e-9b89-eec869672426

Spread prediction.

Hypothesis · H21

4b33f34a-0834-4ef5-b78a-18d6cec760a2

FVT−FOCUS gap prediction.

Premise · P24

f5ada7c4-30c8-405f-8741-cd6fcfb58530

The aggregate evidence the per-domain comparison complements.

Premise · P21

5a078e71-84ff-4039-9e32-7999a5f679f5

The auxiliary embedding space FOCUS needs and the paper does not specify.

Premise · P103

b7e7c5ce-2805-43d7-9f89-ccda038df1b5

The independently built FVT transplant of the full-parameter experiment, whose recipe the FVT initialisation follows and whose sanity loss gate G2 compares with.

Premise · P113

faa76c5f-a834-41c9-8e3f-e4422de4c981

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P117

d9ab6218-cf24-498a-90f1-ebd986fb3508

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P118

ee34e809-5b64-4f37-a761-829e94494bf7

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P119

bbd085f6-467a-47b4-934f-83c089d74d9a

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P120

82fd08a5-cca0-4136-a0c3-a146e9d7d14e

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P122

7ff16749-a07c-49e4-b96e-f230b7c9680c

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P114

f831bb88-fc0d-46ff-88cc-fb9e26e29f8e

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P115

c5737cda-1d82-4f3a-94fa-e13f3310742d

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P116

f81ecf8a-0733-4378-ac05-119e586a18cb

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.