Accepted plan

Sign in with GitHub
← Experiment E5

Immutable accepted plan · retrospective

Matched-budget continued pretraining, banked at 0.5B, with a 1B extension that was not launched

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 2 of 8 July 2026: extend both arms to 1B tokens and anneal both before the headline number. Written after arm B's 100M to 300M scores and citing their gap decay; first committed on 19 July 2026; never launched.

Public source

Plan

Prediction
A substantial share of the change is attributable to continued pretraining; no threshold was fixed.
Protocol
01 streams FineWeb2-HQ and SlimPajama-6B with a fixed shuffle, holds out 2,000 documents of each and counts tokens with both tokenizers; 03 builds the arm-B transplant with Fast Vocabulary Transfer; 02 packs the shared document sequence for each arm; 04 runs a smoke training, then trains arm A and arm B to 0.5B Qwen tokens with milestones every 100M tokens; 05 scores bits per byte on ten texts (first 200 documents each) and 07 scores multiple-choice accuracy for the base, the transplant at time zero and every milestone. The port adds 00, which fetches the APT4 tokenizer files, and 06, which drives 05, 07 and 08 per model. Amendment 1: 08 freezes a digit-probe spec of 300 E2 arithmetic items in three prompt formats and scores every model with greedy 4-shot completions graded by E2's grader; arm-B tokenizers load through a whitespace canary and the frozen APT4 reference. 09 aggregates the rows into the per-split decomposition, the gap per milestone, the probe curves and a paired digit comparison at 500M; 10 writes flat headline values. Amendment 2: both arms are to be extended to 1B tokens from their last training state and then annealed.
Dataset
Training: about 80% FineWeb2-HQ pol_Latn and 20% SlimPajama-6B English by Qwen tokens, one deterministic document sequence for both arms. Evaluation: the first 200 of 2,000 held-out documents of each training source, the first 200 documents of seven E1 corpora, the first 200 math_clean statements, and the first 300 LLMzSzŁ STEM and 250 PES multiple-choice items from E3.
Split
The first 2,000 documents of each seed-42 shuffled source stream are held out and never trained on; training reads the remaining documents in one interleaved order.
Access needs
The APT4 tokenizer is gated on Hugging Face. Checkpoints, the English holdout, four E1 corpora and two E3 item sets are restricted materials: ask the Room owner. The training text and packed data were not kept.
Configurations
Arm A: Qwen2.5-1.5B with its own tokenizer. Arm B: APT4 transplanted with Fast Vocabulary Transfer (piece embeddings as means of the constituent Qwen rows; <s> and </s> from <|endoftext|>; <unk> and 256 byte pieces from the mean embedding; tied embeddings; eos 2). Arm C: untouched. Arms A and B: bf16 full fine-tune, sequence 2048, micro-batch 8, gradient accumulation 32 (524,288 tokens per step), 8-bit AdamW with betas 0.9 and 0.95, weight decay 0.1, clip 1.0, 50 warmup steps then constant learning rate 1e-4, milestones every 1e8 tokens.
Metric
Bits per UTF-8 byte of each text per milestone; gap = arm B minus arm A; continued-pretraining effect = arm A minus arm C. Pooled digit-probe accuracy is the primary downstream mathematics endpoint, multiple-choice accuracy the knowledge endpoint and math_clean bits per byte the compression endpoint.
Seeds
Stream shuffle seed 42 with a 10,000-document buffer; packing and training use no random numbers; one run per arm.
Interpretation rule
None fixed; substantial is not defined. The campaign plan: arms A and B about equal on scientific bits per byte would mean the transplant cost is mostly a continued-pretraining effect. sci_gsm8k is a contamination-decay reference only and is excluded from inference. The matched 0.5B comparison is banked first; the 1B extension with the decay anneal decides between a permanent residual and slow convergence.
Resources
Training on one CUDA GPU, 16 GB, arms serial, about 4 days per arm at 0.5B tokens by the design's estimate; scoring on a separate machine, Apple silicon, 128 GB unified memory.
Prior work
arXiv:2604.10799v1 compares the transplanted models with the untouched originals only; no arm continues the original tokenizer on the same tokens.

Selected exact hypotheses and premises

Hypothesis · H13

f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2

The attribution under test.

Premise · P25

d664a095-2689-445e-ac40-a85c6202a460

The continued pretraining the control arm matches at small scale.

Premise · P12

d5bb75c0-40ed-4d4c-afa4-0fded6fbcb94

The intervention the transplant arm reproduces.

Premise · P22

01087ca0-917c-46f8-a25e-ebc786018d72

The embedding initialisation of the transplant arm.