Immutable accepted plan · prospective
Recovery pretraining of APT4 FVT transplants with 0%, 10% and 30% math and code for 1B tokens at 0.5B, and a 0% against 30% confirmation pair at 1.5B
Amendment 1, committed on 9 July 2026 after the 0.5B analysis refuted H1 and before any confirmation measurement: with no winning mix to confirm, the confirmation re-runs the 0% against 30% contrast at Qwen2.5-1.5B with the same streams, instruments, budget and bins, at micro-batch 8 x gradient accumulation 32.
Public source
https://github.com/stw2/tokenizer-science-tax @ 16a182f7c3592e9e2e84c5a9bc824d2e9b263a06
Reference checked 2026-09-14 13:12 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- At 0.5B as in the locked plan; at 1.5B the same bins apply to conf30 against conf00 with Qwen2.5-1.5B as the base model.
- Protocol
- Per arm on its own GPU, scripts/00_bootstrap_cuda.sh: fetch and check inputs; build the arm's pools (01, gate G6) and pack them by APT4-token shares (02, gate G2); build the FVT transplant twice and compare (E05's 03, gate G3); probe smoke (08, gate G4); resume smoke (E05's 04, gate G5); train 1B tokens (E05's 04) while scoring the base model, the transplant and every 100M-token checkpoint for bits per byte on ten texts (05) and the digit probe (08); per-document scores of the final checkpoint (05b). Per-document scores of the base model (05b); base-model bits per byte on a second backend (G1, 05); analysis (06). Confirmation: the same per-arm pipeline with configs/e6_conf15.json for arms conf00 and conf30 on Qwen2.5-1.5B, per-document scores of Qwen2.5-1.5B, and the same estimators and bins.
- Dataset
- Training: packed streams built from FineWeb2-HQ pol_Latn (revision c0c06e94, inferred), SlimPajama-6B (b5f90f41, inferred), OpenWebMath (open-web-math/open-web-math, no revision recorded or inferable) and codeparrot-clean-train (3e6ab65f, inferred), shuffled with seed 42; the first 2,000 filtered Polish and English documents are skipped as evaluation holdouts; the packed streams, identified by digest, are the training inputs. Evaluation: the first 200 documents of ten texts (Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean) and a 900-completion synthetic digit probe.
- Split
- No train and test split of one dataset: evaluation texts are separate files, and the Polish and English holdouts are skipped in the training streams (gate G6).
- Access needs
- The APT4 tokenizer repository is gated (accept the terms, then use a token); the English evaluation holdout and four evaluation corpora are restricted materials: ask the Room owner, or rebuild them.
- Configurations
- mix00: Polish 80%, English 20%, OpenWebMath 0%, Python code 0%; mix10: Polish 72%, English 18%, OpenWebMath 5%, Python code 5%; mix30: Polish 56%, English 14%, OpenWebMath 15%, Python code 15%; conf00: Polish 80%, English 20%, OpenWebMath 0%, Python code 0%; conf30: Polish 56%, English 14%, OpenWebMath 15%, Python code 15% (shares of APT4 tokens). 0.5B recipe: micro-batch 32 x gradient accumulation 8; 1.5B recipe: micro-batch 8 x gradient accumulation 32; both sequence 2048, learning rate 1e-4, 50 warmup steps, no decay, checkpoints every 100M tokens.
- Metric
- Bits per byte on each text; digit-probe accuracy; rescue fraction R = 1 - (30% arm's gap) / (0% arm's gap) to the base model; pooled Polish bits per byte difference.
- Seeds
- Stream shuffle seed 42; bootstrap seed 20260707; digit-probe item selection seed 20260706 (frozen specification); one training run per arm.
- Interpretation rule
- Per endpoint of H1 (math_clean bits per byte with a paired document bootstrap; digit-probe accuracy with a paired item bootstrap; 2,000 resamples): CONFIRMED if the rescue fraction R is at least 0.8 and its 95% interval's lower bound exceeds 0.5; REFUTED if R is at most 0.5 and the interval's upper bound is below 0.8; VACUOUS if the interval of the 0% arm's gap contains 0; PARTIAL otherwise. H1 is CONFIRMED or REFUTED only when both endpoints agree, else PARTIAL; the design also names REFUTED if the 30% arm is no better than the 0% arm, and a Holm adjustment of H1 and H2, neither computed. H2: CONFIRMED if the 95% interval of the pooled Polish difference (30% minus 0%) is at most 0.05 bits per byte, REFUTED if its lower bound exceeds 0.05, PARTIAL otherwise. H3 is reported without bins. Gates G1 to G6 must pass before results count. Amendment 1: the same rules apply at 1.5B with Qwen2.5-1.5B as the base model.
- Resources
- One 4090-class GPU, 24 GB, per arm; about 8.5 hours per 0.5B arm and an estimated 31 to 46 hours per 1.5B arm; analysis on a CPU.
- Prior work
- arXiv:2604.10799v1 reports GSM8K 85.60 for Bielik-11B-v3.0-Instruct and 80.97 for its APT4 transplant, without the math and code share of its 20B-token adaptation data.
Selected exact hypotheses and premises
Hypothesis · H16
bc2c559b-2e70-44d8-87ff-a07a94962628H3, reported without a decision rule: recovery order.
Hypothesis · H17
8fb0835e-9768-4edc-b293-0d8d150d2fe8H3, reported without a decision rule: dose response of the rescue fraction.
Premise · P25
d664a095-2689-445e-ac40-a85c6202a460The composition of the adaptation data is not given.
Premise · P117
d9ab6218-cf24-498a-90f1-ebd986fb3508Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on the English holdout.
Premise · P118
ee34e809-5b64-4f37-a761-829e94494bf7Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on arXiv abstracts.
Premise · P119
bbd085f6-467a-47b4-934f-83c089d74d9aContinued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on LaTeX method sections.
Premise · P120
82fd08a5-cca0-4136-a0c3-a146e9d7d14eContinued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on Python code.
Premise · P121
78eef7eb-2581-4f13-8cfb-3094c07dbe55Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on math_clean.
Premise · P132
2566f4fe-c7ef-43f0-bb66-0e7ca02bbe3eThe same continued pretraining lowers digit-probe accuracy.
Premise · P108
0c90c839-74ad-4f13-8f54-0b2dfbc9afbbAfter the same continued pretraining, the APT4 transplant trails the original-tokenizer model on the English holdout.
Premise · P109
11728a31-7776-4852-96d3-cb8a28442854After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on arXiv abstracts.
Premise · P110
d3d490b4-d602-49fb-b310-62ebce61bdbdAfter the same continued pretraining, the APT4 transplant trails the original-tokenizer model on LaTeX method sections.
Premise · P111
958beb55-f6ad-48dd-a041-2a5b57024e19After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on Python code.
Premise · P112
1c1e659f-9adf-49a7-8976-649eb293aaa7After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on math_clean.
Premise · P54
2f04e3ac-4c60-437e-93d3-0e5644c55b44APT4 splits every integer into one token per digit.
Premise · P74
b0f7daf6-d1ff-4a89-a255-3c00464a3b96APT4's digit segmentation does not account for the arithmetic accuracy differences of the 11B pair.