Accepted plan

Sign in with GitHub
← Experiment E7

Immutable accepted plan · retrospective

Stage 1 re-score after the tokenizer-load erratum

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 1 of 7 July 2026: the scorer's tokenizer load deleted spaces and newlines for the transplant directories. Every load is now gated on a whitespace round-trip canary with the pristine APT4 tokenizer as fallback, and all four rows are re-scored with unchanged endpoints and bins. Written together with the re-scored rows, so retrospective.

Public source

Plan

Prediction
Unchanged from the first plan version; only measured values and verdicts may move.
Protocol
03, with every tokenizer load gated on a whitespace round-trip canary and the pristine APT4 tokenizer used when a transplant directory loads lossily, re-scores the base model and the three unchanged transplants into a fresh row file; the unchanged 04 recomputes endpoints, bins and gates; 05 writes flat headline values. The auxiliary embeddings and the three builds stand.
Dataset
Qwen2.5-1.5B and APT4; the first 200 documents of Polish web text, English web text, Polish reviews, Polish Wikipedia science articles, Polish PES examination questions, English arXiv abstracts, English LaTeX method sections, English Python code and synthetic arithmetic problems, plus GSM8K problems as a contamination reference; the first 200 MB of Polish FineWeb2-HQ training text for the FOCUS auxiliary embeddings.
Split
Evaluation holdouts and frozen corpora; no transplant is trained in stage 1. GSM8K problems are excluded from every endpoint because the base model has memorised them.
Access needs
The APT4 tokenizer is gated on Hugging Face. The transplant models, the auxiliary corpus, the Polish training text, the English web holdout and four corpora are restricted materials or are rebuilt by script at the pinned revisions.
Configurations
As the first version. Canary: the text 'Oblicz: 37 + 43', a newline and '80' must survive encode and decode; the fallback requires the directory's vocabulary to equal the pristine APT4 tokenizer.json (sha256 536b4dff).
Metric
Bits per byte per domain; init damage, init-quality spread and FVT−FOCUS gap per domain, the last two with 95% paired document bootstrap intervals of 1000 resamples.
Seeds
20260705 for the random initialisation, the fastText training and the bootstrap.
Interpretation rule
H1 passes if the formal-to-prose ratio of mean init damage exceeds 1.15 for both random and FVT initialisation; H2 if FOCUS at most FVT at most random holds on at least 7 of 9 domains; H3 if the formal-to-prose ratio of mean init-quality spread exceeds 1.25; H4 if the mean FVT−FOCUS gap over formal domains exceeds its mean over prose domains. If H1 and H3 both fail, the outcome is a null. Gates: the base model reproduces the full-parameter experiment's time-zero base row within 0.005 bits per byte on 9 domains; the rebuilt FVT transplant's sanity loss is about 10.3; each initialisation rebuilds byte-identically; FVT and FOCUS are equal on overlap rows; vocabulary 32,000 with tied embeddings. Acceptance of the re-score: the re-scored base row equals the earlier base row, the re-scored FVT row equals the independently built FVT transplant's time-zero row on every shared domain, and G1 still passes.
Resources
Apple silicon, 128 GB unified memory: MPS, about one hour of scoring.
Prior work
arXiv:2604.10799v1 chooses FOCUS citing aggregate results on Bielik 1.5B v3 and reports no per-domain comparison of initialisations.

Selected exact hypotheses and premises

Hypothesis · H18

c8bc4afe-d6ab-4475-b668-4f01d4d149b2

Formal-damage prediction.

Hypothesis · H19

2174054b-06f0-40a7-b2f7-96cecb0b7c75

Ranking prediction.

Hypothesis · H20

3d170e90-be69-449e-9b89-eec869672426

Spread prediction.

Hypothesis · H21

4b33f34a-0834-4ef5-b78a-18d6cec760a2

FVT−FOCUS gap prediction.

Premise · P24

f5ada7c4-30c8-405f-8741-cd6fcfb58530

The aggregate evidence the per-domain comparison complements.

Premise · P103

b7e7c5ce-2805-43d7-9f89-ccda038df1b5

The independently built FVT transplant of the full-parameter experiment, whose time-zero bits-per-byte rows acceptance gate (ii) of the re-score compares with.

Premise · P113

faa76c5f-a834-41c9-8e3f-e4422de4c981

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P117

d9ab6218-cf24-498a-90f1-ebd986fb3508

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P118

ee34e809-5b64-4f37-a761-829e94494bf7

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P119

bbd085f6-467a-47b4-934f-83c089d74d9a

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P120

82fd08a5-cca0-4136-a0c3-a146e9d7d14e

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P122

7ff16749-a07c-49e4-b96e-f230b7c9680c

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P114

f831bb88-fc0d-46ff-88cc-fb9e26e29f8e

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P115

c5737cda-1d82-4f3a-94fa-e13f3310742d

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.

Premise · P116

f81ecf8a-0733-4378-ac05-119e586a18cb

Its starting value is the untouched base model's bits per byte on this text, the target gate G1 reproduces.