Accepted plan

Sign in with GitHub
← Experiment E6

Immutable accepted plan · prospective

Recovery pretraining of APT4 FVT transplants with 0%, 10% and 30% math and code for 1B tokens at 0.5B, and a 0% against 30% confirmation pair at 1.5B

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 1, committed on 9 July 2026 after the 0.5B analysis refuted H1 and before any confirmation measurement: with no winning mix to confirm, the confirmation re-runs the 0% against 30% contrast at Qwen2.5-1.5B with the same streams, instruments, budget and bins, at micro-batch 8 x gradient accumulation 32.

Public source

Plan

Prediction
At 0.5B as in the locked plan; at 1.5B the same bins apply to conf30 against conf00 with Qwen2.5-1.5B as the base model.
Protocol
Per arm on its own GPU, scripts/00_bootstrap_cuda.sh: fetch and check inputs; build the arm's pools (01, gate G6) and pack them by APT4-token shares (02, gate G2); build the FVT transplant twice and compare (E05's 03, gate G3); probe smoke (08, gate G4); resume smoke (E05's 04, gate G5); train 1B tokens (E05's 04) while scoring the base model, the transplant and every 100M-token checkpoint for bits per byte on ten texts (05) and the digit probe (08); per-document scores of the final checkpoint (05b). Per-document scores of the base model (05b); base-model bits per byte on a second backend (G1, 05); analysis (06). Confirmation: the same per-arm pipeline with configs/e6_conf15.json for arms conf00 and conf30 on Qwen2.5-1.5B, per-document scores of Qwen2.5-1.5B, and the same estimators and bins.
Dataset
Training: packed streams built from FineWeb2-HQ pol_Latn (revision c0c06e94, inferred), SlimPajama-6B (b5f90f41, inferred), OpenWebMath (open-web-math/open-web-math, no revision recorded or inferable) and codeparrot-clean-train (3e6ab65f, inferred), shuffled with seed 42; the first 2,000 filtered Polish and English documents are skipped as evaluation holdouts; the packed streams, identified by digest, are the training inputs. Evaluation: the first 200 documents of ten texts (Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean) and a 900-completion synthetic digit probe.
Split
No train and test split of one dataset: evaluation texts are separate files, and the Polish and English holdouts are skipped in the training streams (gate G6).
Access needs
The APT4 tokenizer repository is gated (accept the terms, then use a token); the English evaluation holdout and four evaluation corpora are restricted materials: ask the Room owner, or rebuild them.
Configurations
mix00: Polish 80%, English 20%, OpenWebMath 0%, Python code 0%; mix10: Polish 72%, English 18%, OpenWebMath 5%, Python code 5%; mix30: Polish 56%, English 14%, OpenWebMath 15%, Python code 15%; conf00: Polish 80%, English 20%, OpenWebMath 0%, Python code 0%; conf30: Polish 56%, English 14%, OpenWebMath 15%, Python code 15% (shares of APT4 tokens). 0.5B recipe: micro-batch 32 x gradient accumulation 8; 1.5B recipe: micro-batch 8 x gradient accumulation 32; both sequence 2048, learning rate 1e-4, 50 warmup steps, no decay, checkpoints every 100M tokens.
Metric
Bits per byte on each text; digit-probe accuracy; rescue fraction R = 1 - (30% arm's gap) / (0% arm's gap) to the base model; pooled Polish bits per byte difference.
Seeds
Stream shuffle seed 42; bootstrap seed 20260707; digit-probe item selection seed 20260706 (frozen specification); one training run per arm.
Interpretation rule
Per endpoint of H1 (math_clean bits per byte with a paired document bootstrap; digit-probe accuracy with a paired item bootstrap; 2,000 resamples): CONFIRMED if the rescue fraction R is at least 0.8 and its 95% interval's lower bound exceeds 0.5; REFUTED if R is at most 0.5 and the interval's upper bound is below 0.8; VACUOUS if the interval of the 0% arm's gap contains 0; PARTIAL otherwise. H1 is CONFIRMED or REFUTED only when both endpoints agree, else PARTIAL; the design also names REFUTED if the 30% arm is no better than the 0% arm, and a Holm adjustment of H1 and H2, neither computed. H2: CONFIRMED if the 95% interval of the pooled Polish difference (30% minus 0%) is at most 0.05 bits per byte, REFUTED if its lower bound exceeds 0.05, PARTIAL otherwise. H3 is reported without bins. Gates G1 to G6 must pass before results count. Amendment 1: the same rules apply at 1.5B with Qwen2.5-1.5B as the base model.
Resources
One 4090-class GPU, 24 GB, per arm; about 8.5 hours per 0.5B arm and an estimated 31 to 46 hours per 1.5B arm; analysis on a CPU.
Prior work
arXiv:2604.10799v1 reports GSM8K 85.60 for Bielik-11B-v3.0-Instruct and 80.97 for its APT4 transplant, without the math and code share of its 20B-token adaptation data.

Selected exact hypotheses and premises

Hypothesis · H14

f89e740e-7204-4148-8934-85fbc73f323c

H1, primary: the rescue.

Hypothesis · H15

ed4b79d2-8435-4a8b-b3c2-f466d5084d7b

H2: the Polish cost of the rescue.

Hypothesis · H16

bc2c559b-2e70-44d8-87ff-a07a94962628

H3, reported without a decision rule: recovery order.

Hypothesis · H17

8fb0835e-9768-4edc-b293-0d8d150d2fe8

H3, reported without a decision rule: dose response of the rescue fraction.

Premise · P25

d664a095-2689-445e-ac40-a85c6202a460

The composition of the adaptation data is not given.

Premise · P29

36da68d9-97cf-4575-ad42-ed56e0e895e9

The GSM8K regression under test.

Premise · P117

d9ab6218-cf24-498a-90f1-ebd986fb3508

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on the English holdout.

Premise · P118

ee34e809-5b64-4f37-a761-829e94494bf7

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on arXiv abstracts.

Premise · P119

bbd085f6-467a-47b4-934f-83c089d74d9a

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on LaTeX method sections.

Premise · P120

82fd08a5-cca0-4136-a0c3-a146e9d7d14e

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on Python code.

Premise · P121

78eef7eb-2581-4f13-8cfb-3094c07dbe55

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on math_clean.

Premise · P132

2566f4fe-c7ef-43f0-bb66-0e7ca02bbe3e

The same continued pretraining lowers digit-probe accuracy.

Premise · P108

0c90c839-74ad-4f13-8f54-0b2dfbc9afbb

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on the English holdout.

Premise · P109

11728a31-7776-4852-96d3-cb8a28442854

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on arXiv abstracts.

Premise · P110

d3d490b4-d602-49fb-b310-62ebce61bdbd

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on LaTeX method sections.

Premise · P111

958beb55-f6ad-48dd-a041-2a5b57024e19

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on Python code.

Premise · P112

1c1e659f-9adf-49a7-8976-649eb293aaa7

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on math_clean.

Premise · P133

683a738d-67cc-4630-b105-e5ad71c2d77a

The transplant also trails on the digit probe.

Premise · P54

2f04e3ac-4c60-437e-93d3-0e5644c55b44

APT4 splits every integer into one token per digit.

Premise · P74

b0f7daf6-d1ff-4a89-a255-3c00464a3b96

APT4's digit segmentation does not account for the arithmetic accuracy differences of the 11B pair.