Immutable accepted plan · prospective
Stage 2: 150M tokens of embeddings-only continued pretraining of the three transplants, with bits-per-byte recovery at six checkpoints
Amendment 2 of 9 July 2026: locks the stage-2 recipe, milestones, bins and gates before any stage-2 training. After the erratum the FVT−FOCUS gap is near zero on two domains, so the primary endpoint becomes the random-minus-FVT gap on formal domains and FOCUS−FVT persistence becomes secondary on the seven domains whose time-zero interval excludes zero.
Public source
https://github.com/stw2/tokenizer-science-tax @ 80716b075735d909dacf90838f0e4ced6af46be8
Reference checked 2026-09-14 13:12 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- No direction was pre-registered: the primary endpoint classifies the random-minus-FVT gap on each formal domain at 150M tokens as CLOSED, PARTIAL or RESIDUAL, and the secondary endpoint classifies the persistence of the FVT−FOCUS gap as VANISHED, SHRINKS or PERSISTS.
- Protocol
- drive_stage2.sh runs 07 serially for FVT, FOCUS and random. Training updates only the tied input embedding matrix on the packed APT4 token stream read from offset 0, with 8-bit AdamW, learning rate 1e-4, 50 warmup steps and no decay, weight decay 0.1, clipping 1.0, sequence 2,048, micro-batch 8 and gradient accumulation 32 (524,288 tokens per step), bf16, non-reentrant gradient checkpointing and fused cross-entropy, and saves checkpoints at 30M, 60M, 90M, 100M, 120M and 150M tokens. The unchanged 03 scores each checkpoint on the ten domains; 09 computes the stage-2 endpoints with paired document bootstrap intervals; 05 writes flat headline values. Time zero is the re-scored stage-1 rows.
- Dataset
- Training: the first 150,470,656 tokens of the packed APT4 token stream of Polish FineWeb2-HQ and English SlimPajama-6B text built for full-parameter continued pretraining of the FVT transplant. Evaluation: the ten stage-1 domains, first 200 documents each.
- Split
- The evaluation holdouts are excluded from the training stream by its builder; the evaluation domains are unchanged from stage 1.
- Access needs
- The packed token stream and the stage-2 checkpoints are restricted; the stream's digest was never recorded and it is rebuilt by script. Training needs a CUDA GPU of about 16 GB; the APT4 tokenizer is gated.
- Configurations
- Arms FVT, FOCUS, random in that order; trainable set model.embed_tokens.weight only, output head tied; milestones 30M, 60M, 90M, 100M, 120M, 150M tokens; checkpoint tokenizers are read through the scorer's whitespace-gated loader.
- Metric
- Primary: random-minus-FVT bits per byte at 150M tokens on each formal domain with a 95% paired bootstrap interval. Secondary: FVT−FOCUS bits per byte at 150M tokens divided by its time-zero value on seven domains. Descriptive: recovery curves, tokens to 90% of each arm's own 150M-token recovery, and the FVT arm's recovery at 100M tokens divided by the recovery of full-parameter training of the same transplant on the same stream.
- Seeds
- 20260705 for the bootstrap; training reads the stream sequentially.
- Interpretation rule
- Per formal domain at 150M tokens: CLOSED if the interval includes 0 or the gap is below 0.05 bits per byte; RESIDUAL if the gap is at least 0.20 with the interval excluding 0; PARTIAL otherwise. Class: HEALS if all three are CLOSED, FLOOR-RESIDUAL if any is RESIDUAL, PARTIAL otherwise. Per eligible domain, with P the 150M-token gap divided by the time-zero gap, VANISHED if the 150M interval includes 0 or P is at most 0.25; PERSISTS if P is at least 0.75 with the interval excluding 0; SHRINKS otherwise. Gates: identical model digests on the building and the training machine; trainable-set audit at start; body frozen and embeddings moved after a smoke run; resume continuity; the scorer's fallback to the pristine APT4 tokenizer on the first scored checkpoint. Halt if the gradient norm stays near zero or smoke throughput is below 2,000 tokens per second.
- Resources
- Consumer CUDA GPU, 16 GB, about 9 hours per arm; Apple silicon, 128 GB unified memory (MPS) for scoring, about 15 minutes per checkpoint.
- Prior work
- arXiv:2604.10799v1 describes 4B tokens of continued pretraining that update the input embeddings, the output head and four boundary layers before full adaptation, and compares no initialisations under that stage.
Selected exact hypotheses and premises
Hypothesis · H18
c8bc4afe-d6ab-4475-b668-4f01d4d149b2Stage 2 follows the random-initialisation damage on formal domains under embeddings-only training.
Hypothesis · H20
3d170e90-be69-449e-9b89-eec869672426Stage 2 follows the advantage of informed initialisation over random on formal domains.
Hypothesis · H21
4b33f34a-0834-4ef5-b78a-18d6cec760a2Stage 2 follows the FVT−FOCUS gap under embeddings-only training.
Premise · P26
4aeee3b3-fb89-4ae6-9e40-9668ffa62e84The embeddings-first stage of vocabulary adaptation that stage 2 mirrors, here with only the embeddings trainable.
Premise · P23
cae5a2ea-f055-4e01-bb4b-86bb0b6c0fd3Random initialisation is said to relearn embeddings from scratch and converge slowly.
Premise · P103
b7e7c5ce-2805-43d7-9f89-ccda038df1b5The full-parameter experiment's FVT transplant, whose training recipe stage 2 repeats with only the embeddings trainable.
Premise · P124
131e534a-1b09-40c1-95fc-a4d5d63faae0Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.
Premise · P125
f1355715-5d7f-4a04-a4bf-249f33aa288eFull-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.
Premise · P126
b11964d8-81ba-4bf6-9364-302e0e5a677cFull-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.
Premise · P127
ebd85d87-d358-44c8-aa85-5e8246d99b2fFull-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.
Premise · P128
4a7dff3a-4d95-465e-a43e-bb0df933fc9eFull-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.
Premise · P129
e2da0e41-7694-4b22-a1ea-14d13dc51f38Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.