Accepted plan

Sign in with GitHub
← Experiment E7

Immutable accepted plan · prospective

Stage 2: 150M tokens of embeddings-only continued pretraining of the three transplants, with bits-per-byte recovery at six checkpoints

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 2 of 9 July 2026: locks the stage-2 recipe, milestones, bins and gates before any stage-2 training. After the erratum the FVT−FOCUS gap is near zero on two domains, so the primary endpoint becomes the random-minus-FVT gap on formal domains and FOCUS−FVT persistence becomes secondary on the seven domains whose time-zero interval excludes zero.

Public source

Plan

Prediction
No direction was pre-registered: the primary endpoint classifies the random-minus-FVT gap on each formal domain at 150M tokens as CLOSED, PARTIAL or RESIDUAL, and the secondary endpoint classifies the persistence of the FVT−FOCUS gap as VANISHED, SHRINKS or PERSISTS.
Protocol
drive_stage2.sh runs 07 serially for FVT, FOCUS and random. Training updates only the tied input embedding matrix on the packed APT4 token stream read from offset 0, with 8-bit AdamW, learning rate 1e-4, 50 warmup steps and no decay, weight decay 0.1, clipping 1.0, sequence 2,048, micro-batch 8 and gradient accumulation 32 (524,288 tokens per step), bf16, non-reentrant gradient checkpointing and fused cross-entropy, and saves checkpoints at 30M, 60M, 90M, 100M, 120M and 150M tokens. The unchanged 03 scores each checkpoint on the ten domains; 09 computes the stage-2 endpoints with paired document bootstrap intervals; 05 writes flat headline values. Time zero is the re-scored stage-1 rows.
Dataset
Training: the first 150,470,656 tokens of the packed APT4 token stream of Polish FineWeb2-HQ and English SlimPajama-6B text built for full-parameter continued pretraining of the FVT transplant. Evaluation: the ten stage-1 domains, first 200 documents each.
Split
The evaluation holdouts are excluded from the training stream by its builder; the evaluation domains are unchanged from stage 1.
Access needs
The packed token stream and the stage-2 checkpoints are restricted; the stream's digest was never recorded and it is rebuilt by script. Training needs a CUDA GPU of about 16 GB; the APT4 tokenizer is gated.
Configurations
Arms FVT, FOCUS, random in that order; trainable set model.embed_tokens.weight only, output head tied; milestones 30M, 60M, 90M, 100M, 120M, 150M tokens; checkpoint tokenizers are read through the scorer's whitespace-gated loader.
Metric
Primary: random-minus-FVT bits per byte at 150M tokens on each formal domain with a 95% paired bootstrap interval. Secondary: FVT−FOCUS bits per byte at 150M tokens divided by its time-zero value on seven domains. Descriptive: recovery curves, tokens to 90% of each arm's own 150M-token recovery, and the FVT arm's recovery at 100M tokens divided by the recovery of full-parameter training of the same transplant on the same stream.
Seeds
20260705 for the bootstrap; training reads the stream sequentially.
Interpretation rule
Per formal domain at 150M tokens: CLOSED if the interval includes 0 or the gap is below 0.05 bits per byte; RESIDUAL if the gap is at least 0.20 with the interval excluding 0; PARTIAL otherwise. Class: HEALS if all three are CLOSED, FLOOR-RESIDUAL if any is RESIDUAL, PARTIAL otherwise. Per eligible domain, with P the 150M-token gap divided by the time-zero gap, VANISHED if the 150M interval includes 0 or P is at most 0.25; PERSISTS if P is at least 0.75 with the interval excluding 0; SHRINKS otherwise. Gates: identical model digests on the building and the training machine; trainable-set audit at start; body frozen and embeddings moved after a smoke run; resume continuity; the scorer's fallback to the pristine APT4 tokenizer on the first scored checkpoint. Halt if the gradient norm stays near zero or smoke throughput is below 2,000 tokens per second.
Resources
Consumer CUDA GPU, 16 GB, about 9 hours per arm; Apple silicon, 128 GB unified memory (MPS) for scoring, about 15 minutes per checkpoint.
Prior work
arXiv:2604.10799v1 describes 4B tokens of continued pretraining that update the input embeddings, the output head and four boundary layers before full adaptation, and compares no initialisations under that stage.

Selected exact hypotheses and premises

Hypothesis · H18

c8bc4afe-d6ab-4475-b668-4f01d4d149b2

Stage 2 follows the random-initialisation damage on formal domains under embeddings-only training.

Hypothesis · H20

3d170e90-be69-449e-9b89-eec869672426

Stage 2 follows the advantage of informed initialisation over random on formal domains.

Hypothesis · H21

4b33f34a-0834-4ef5-b78a-18d6cec760a2

Stage 2 follows the FVT−FOCUS gap under embeddings-only training.

Premise · P26

4aeee3b3-fb89-4ae6-9e40-9668ffa62e84

The embeddings-first stage of vocabulary adaptation that stage 2 mirrors, here with only the embeddings trainable.

Premise · P23

cae5a2ea-f055-4e01-bb4b-86bb0b6c0fd3

Random initialisation is said to relearn embeddings from scratch and converge slowly.

Premise · P103

b7e7c5ce-2805-43d7-9f89-ccda038df1b5

The full-parameter experiment's FVT transplant, whose training recipe stage 2 repeats with only the embeddings trainable.

Premise · P124

131e534a-1b09-40c1-95fc-a4d5d63faae0

Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.

Premise · P125

f1355715-5d7f-4a04-a4bf-249f33aa288e

Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.

Premise · P126

b11964d8-81ba-4bf6-9364-302e0e5a677c

Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.

Premise · P127

ebd85d87-d358-44c8-aa85-5e8246d99b2f

Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.

Premise · P128

4a7dff3a-4d95-465e-a43e-bb0df933fc9e

Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.

Premise · P129

e2da0e41-7694-4b22-a1ea-14d13dc51f38

Full-parameter continued pretraining of the same FVT transplant on the same token stream, whose time-zero and 100M-token rows give the denominator of the embeddings-only share of recovery.