Immutable accepted plan · retrospective
Five 32k SentencePiece BPE tokenizers differing in training-data allocation, audited for digit and separator handling and measured on a fertility atlas of seven corpora
Amendment 2, 5 July 2026, written after the determinism gate had failed: the gate compares model content with the stored output path cleared, because the retrained model differed from the first only in that field while its pieces and tokenizer.json were identical.
Public source
https://github.com/stw2/tokenizer-science-tax @ 7a2cffbe78cf5f6b8ee85b8ea6dd34fcb366b296
Reference checked 2026-09-14 13:15 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- The science-slice arm's pooled English-science excess tax is at most half the Polish-only arm's at a Polish token cost below 5% pooled and 7% per corpus; the separator arm encodes each of the three characters in one token between digits at a pooled Polish token cost below 0.5%.
- Protocol
- 01 builds five frozen training texts from pinned dataset revisions (Polish web text, arXiv LaTeX, Python code, mathematical web text, general English web text) with byte quotas and no randomness; 02 assembles each arm's input from line-bounded prefixes, trains SentencePiece BPE at 32,000 pieces, converts to a fast tokenizer in APT4's pipeline shape and runs gates G1, G4 and G5 and, retraining the Polish-only arm, the determinism gate G3; 03 audits digit and separator tokenization on synthetic number renderings (gate G2); 04 measures tokens per word, characters per token and the tax against the Mistral-derived tokenizer for fourteen tokenizers on seven corpora and two preamble texts with a paired document bootstrap and checks that the nine reference tokenizers' values reproduce exactly; 05 computes the endpoints and bins; 06 writes flat headline values; 07 writes the content manifest.
- Dataset
- Training: 1 GiB of text per arm from FineWeb2-HQ pol_Latn, proof-pile-2 arXiv train shard 0, proof-pile-2 OpenWebMath train shard 0, codeparrot-clean-train and SlimPajama-6B at pinned revisions. Evaluation: seven corpora of about 120k words each (English arXiv abstracts, LaTeX method sections, Python code, GSM8K problems; Polish PES questions, Wikipedia science articles, reviews) and the Polish and English preamble protocol texts; synthetic integer, big-integer, rendering and sign sets for the digit audit.
- Split
- Training text is kept apart from the evaluation corpora by source split, repository and a Wikipedia URL filter, and GSM8K is excluded; the first 2,000 documents of the Polish and English streams are skipped. Every evaluation document is measured.
- Access needs
- Training text, arm inputs and trainer work files are restricted materials that rebuild from pinned dataset revisions; the reference tokenizers include gated repositories and four evaluation corpora are restricted.
- Configurations
- Arms: Polish-only (100% Polish); Polish-only with separator pieces (the same text); science slice (80% Polish, 20% science in equal thirds arXiv LaTeX, Python code, mathematical web text, with separator pieces); science slice with whitespace-only pieces allowed (the same text); balanced (60% Polish, 20% general English, 20% science, with separator pieces). Polish portions are nested prefixes of one stream and the science slices are identical across arms; 1 GiB of text per arm. Shared: BPE, vocabulary 32,000, character coverage 0.9995, identity normalisation, dummy prefix, byte fallback, digits split, maximum piece length 16, no input shuffling, APT4's special tokens. Separator pieces: U+00A0, U+202F and U+2212 added as user-defined symbols at ids 263 to 265.
- Metric
- Pooled excess tax: tokens over the four English corpora divided by the Mistral-derived tokenizer's tokens, minus 1. R: an arm's pooled excess tax divided by the Polish-only arm's. Polish cost: an arm's tokens over the three Polish corpora divided by the Polish-only arm's, minus 1. Separator tokens: tokens a character adds between digits. 95% percentile intervals over 1000 stratified paired document resamples.
- Seeds
- 20260703 (bootstrap); corpus assembly and tokenizer training use no randomness
- Interpretation rule
- H1 on the science-slice arm: ratio of pooled excess taxes at most 0.50 confirmed, above 0.50 up to 0.75 partial, above 0.75 refuted; the whitespace arm is secondary with the same bins. H2 on the science-slice arm: pooled Polish cost below 0.05 and every per-corpus cost below 0.07 confirmed, pooled cost from 0.05 to below 0.08 partial, 0.08 or more refuted. Allocation beats size is confirmed if H1 and H2 are both confirmed for the science-slice arm, partial if either is partial or only the whitespace arm confirms both, refuted otherwise. H3 on the separator arm: each character one token between digits and pooled Polish cost below 0.005 confirmed, one of the two partial, neither refuted. The balanced arm is exploratory without a bin. Gates G1 (APT4 pipeline shape and id layout), G2 (single-digit policy), G4 (SentencePiece and fast tokenizer agree on 10,000 lines) and G5 (separator pieces present exactly in the arms that add them) must pass before the atlas is read; G3 requires identical model content with trainer_spec.model_prefix cleared, a byte-identical tokenizer.json and a byte-identical atlas when the Polish-only arm is trained twice.
- Resources
- CPU only; about one hour in the original run; about 7 GB of disk for rebuilt training text and trainer files.
- Prior work
- arXiv:2604.10799v1 holds APT4 at about 32k tokens and gives no training corpus, allocation or digit and punctuation policy for it. The fertility atlas of the same corpora measured APT4's tax against the Mistral-derived tokenizer.
Selected exact hypotheses and premises
Hypothesis · H24
d1a46254-542b-43a9-b64b-3536a37e2efdH1, the headline prediction on the science-slice arm.
Hypothesis · H26
6b80001d-30ca-46bf-8aa3-74928a619665H3, the separator pieces on the Polish-only arm.
Premise · P52
3d1e1910-f471-4b1a-be65-5e2db901f9a0APT4's English fertility tax depends on the science domain.
Premise · P15
e6daa427-c8d9-4276-bc5b-aacb1fc82344APT4's vocabulary size is held at about 32k tokens.
Premise · P20
85dfefb0-eb84-4deb-b93c-7b494a10b283Digit and special-character handling can change token efficiency.