Accepted plan

Sign in with GitHub
← Experiment E9

Immutable accepted plan · retrospective

Five 32k SentencePiece BPE tokenizers differing in training-data allocation, audited for digit and separator handling and measured on a fertility atlas of seven corpora

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 1, 5 July 2026, before any arm trained: the separator pieces are added as user-defined symbols instead of required characters, because the trainer aborts when a required character does not occur in the training text; user-defined symbols also act as boundaries no merge crosses.

Public source

Plan

Prediction
The science-slice arm's pooled English-science excess tax is at most half the Polish-only arm's at a Polish token cost below 5% pooled and 7% per corpus; the separator arm encodes each of the three characters in one token between digits at a pooled Polish token cost below 0.5%.
Protocol
01 builds five frozen training texts from pinned dataset revisions (Polish web text, arXiv LaTeX, Python code, mathematical web text, general English web text) with byte quotas and no randomness; 02 assembles each arm's input from line-bounded prefixes, trains SentencePiece BPE at 32,000 pieces, converts to a fast tokenizer in APT4's pipeline shape and runs gates G1, G4 and G5 and, retraining the Polish-only arm, the determinism gate G3; 03 audits digit and separator tokenization on synthetic number renderings (gate G2); 04 measures tokens per word, characters per token and the tax against the Mistral-derived tokenizer for fourteen tokenizers on seven corpora and two preamble texts with a paired document bootstrap and checks that the nine reference tokenizers' values reproduce exactly; 05 computes the endpoints and bins; 06 writes flat headline values; 07 writes the content manifest.
Dataset
Training: 1 GiB of text per arm from FineWeb2-HQ pol_Latn, proof-pile-2 arXiv train shard 0, proof-pile-2 OpenWebMath train shard 0, codeparrot-clean-train and SlimPajama-6B at pinned revisions. Evaluation: seven corpora of about 120k words each (English arXiv abstracts, LaTeX method sections, Python code, GSM8K problems; Polish PES questions, Wikipedia science articles, reviews) and the Polish and English preamble protocol texts; synthetic integer, big-integer, rendering and sign sets for the digit audit.
Split
Training text is kept apart from the evaluation corpora by source split, repository and a Wikipedia URL filter, and GSM8K is excluded; the first 2,000 documents of the Polish and English streams are skipped. Every evaluation document is measured.
Access needs
Training text, arm inputs and trainer work files are restricted materials that rebuild from pinned dataset revisions; the reference tokenizers include gated repositories and four evaluation corpora are restricted.
Configurations
Arms: Polish-only (100% Polish); Polish-only with separator pieces (the same text); science slice (80% Polish, 20% science in equal thirds arXiv LaTeX, Python code, mathematical web text, with separator pieces); science slice with whitespace-only pieces allowed (the same text); balanced (60% Polish, 20% general English, 20% science, with separator pieces). Polish portions are nested prefixes of one stream and the science slices are identical across arms; 1 GiB of text per arm. Shared: BPE, vocabulary 32,000, character coverage 0.9995, identity normalisation, dummy prefix, byte fallback, digits split, maximum piece length 16, no input shuffling, APT4's special tokens. Separator pieces: U+00A0, U+202F and U+2212 added as user-defined symbols at ids 263 to 265.
Metric
Pooled excess tax: tokens over the four English corpora divided by the Mistral-derived tokenizer's tokens, minus 1. R: an arm's pooled excess tax divided by the Polish-only arm's. Polish cost: an arm's tokens over the three Polish corpora divided by the Polish-only arm's, minus 1. Separator tokens: tokens a character adds between digits. 95% percentile intervals over 1000 stratified paired document resamples.
Seeds
20260703 (bootstrap); corpus assembly and tokenizer training use no randomness
Interpretation rule
H1 on the science-slice arm: ratio of pooled excess taxes at most 0.50 confirmed, above 0.50 up to 0.75 partial, above 0.75 refuted; the whitespace arm is secondary with the same bins. H2 on the science-slice arm: pooled Polish cost below 0.05 and every per-corpus cost below 0.07 confirmed, pooled cost from 0.05 to below 0.08 partial, 0.08 or more refuted. Allocation beats size is confirmed if H1 and H2 are both confirmed for the science-slice arm, partial if either is partial or only the whitespace arm confirms both, refuted otherwise. H3 on the separator arm: each character one token between digits and pooled Polish cost below 0.005 confirmed, one of the two partial, neither refuted. The balanced arm is exploratory without a bin. Gates G1 (APT4 pipeline shape and id layout), G2 (single-digit policy), G4 (SentencePiece and fast tokenizer agree on 10,000 lines) and G5 (separator pieces present exactly in the arms that add them) must pass before the atlas is read; G3 requires a byte-identical model file, a byte-identical tokenizer.json and a byte-identical atlas when the Polish-only arm is trained twice.
Resources
CPU only; about one hour in the original run; about 7 GB of disk for rebuilt training text and trainer files.
Prior work
arXiv:2604.10799v1 holds APT4 at about 32k tokens and gives no training corpus, allocation or digit and punctuation policy for it. The fertility atlas of the same corpora measured APT4's tax against the Mistral-derived tokenizer.

Selected exact hypotheses and premises

Hypothesis · H24

d1a46254-542b-43a9-b64b-3536a37e2efd

H1, the headline prediction on the science-slice arm.

Hypothesis · H25

2be4b5ff-86ee-4e87-b3a4-77932559f05f

H2, the Polish cost of the science slice.

Hypothesis · H26

6b80001d-30ca-46bf-8aa3-74928a619665

H3, the separator pieces on the Polish-only arm.

Premise · P52

3d1e1910-f471-4b1a-be65-5e2db901f9a0

APT4's English fertility tax depends on the science domain.

Premise · P15

e6daa427-c8d9-4276-bc5b-aacb1fc82344

APT4's vocabulary size is held at about 32k tokens.

Premise · P20

85dfefb0-eb84-4deb-b93c-7b494a10b283

Digit and special-character handling can change token efficiency.