Immutable accepted plan · retrospective
Digit and separator policy audit of nine tokenizers, answer-only arithmetic probes on the 11B Bielik pair in four number formats and two instruction languages, and a 250-problem GSM8K subsample
The plan as run on 3 July 2026, written from a plan file kept outside the repository and cited by the report, and from the campaign plan's E2 section; no design was committed before the measurements, and scripts, results and report were committed together.
Public source
https://github.com/stw2/tokenizer-science-tax @ c43f19850803969bc1fefe537caf14e54cfd9a9f
Reference checked 2026-09-14 13:08 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- APT4 segments digits one token per digit like the Mistral-derived tokenizer, APT3 by free BPE with digit runs of at least four, and the other audited tokenizers as their published policies state; APT4 needs more tokens per number with no-break-space grouping; Bielik-PL-11B-v3.0-Instruct's accuracy deficit is larger with comma-grouped than with space-grouped numbers; replacing space grouping with no-break-space grouping lowers its accuracy more than Bielik-11B-v3.0-Instruct's; separators, signs and units are segmented identically except the no-break space and the U+2212 minus sign.
- Protocol
- 01 downloads both 11B models at pinned revisions and checks that their tokenizer.json files equal E1's; 02 converts both to MLX bf16 with the same mlx-lm version; 03 generates the number sets S1 to S8 and the GSM8K numeral inventory S9; 04 classifies the digit policy of nine tokenizers, measures tokens per number in six renderings and tabulates separator, sign and unit segmentation, counting tokens that overlap the number after the prefix "a " without special tokens; 05 generates 500 items per task for addition, subtraction, comparison, sorting and unit conversion, each rendered in four number formats and two instruction languages; 06 runs a 50-item smoke gate (one beginning-of-sequence token, byte-identical greedy repeat, throughput) and then all 20,000 generations per model with greedy decoding and at most 256 tokens; 07 grades with a dual-locale grader gated by golden tests and computes per-cell accuracy, the contrasts C1, C4', C2 and C5 with paired bootstrap intervals and Holm adjustment, and carry gradients; 08 runs 250 GSM8K test problems with five fixed examples and greedy decoding at most 512 tokens; 09 writes flat headline values.
- Dataset
- Synthetic number sets S1 to S8 and 2,500 arithmetic items generated with seed 20260703; GSM8K main configuration: 250 test problems, five training problems as examples, and both splits for the numeral inventory.
- Split
- Number sets and probe items are generated, not split. GSM8K: 250 test problems stratified by answer digit length; five training problems drawn once as fixed examples.
- Access needs
- The two 11B model repositories are gated on Hugging Face (accept the terms, then use a token); the downloaded snapshots and MLX conversions are restricted materials. Probes and GSM8K generation need Apple silicon with MLX and about 89 GB of disk; E1's tokenizers are rebuilt by E1's fetch script.
- Configurations
- Audit renderings: bare, comma grouping, space grouping, no-break-space grouping, narrow no-break-space grouping, period grouping. Probe cells: bare, comma-grouped, space-grouped and no-break-space-grouped numbers crossed with English and Polish instructions. Tasks: addition and subtraction of 2 to 7 digit operands, comparison with decimal traps, sorting of five numbers, unit conversion by powers of ten. Greedy decoding; 256 generated tokens for probes, 512 for GSM8K.
- Metric
- Digit policy class and tokens per number per rendering (audit); share of probe items with the correct value; contrasts C1 (model by format interaction, comma against space grouping), C4' (no-break-space differential), C2 (within-model format effects) and C5 (model main effect) on informative strata and items of at least four digits, with 95% paired bootstrap intervals and Holm adjustment over C1, C4', C2 for the PL model and C5; GSM8K accuracy and paired difference.
- Seeds
- 20260703 (number sets, probe items, GSM8K subsample and examples, bootstrap)
- Interpretation rule
- H1 is killed by any systematic segmentation difference between APT4 and the Mistral-derived tokenizer on bare or comma-grouped numerals beyond special-character handling. H2 is killed at a rendering where the ratio of tokens per number differs by less than 5%. H3 is killed if the pooled interaction interval covers 0 or its sign is opposite. H4 is killed if the interval of the no-break-space differential covers 0. H5 is killed by identical segmentation everywhere. Tokenizer claims come only from format-differential contrasts; model main effects are attributed to the transplant pipeline. Mechanism: refutes if digit segmentation is identical on bare and comma-grouped numerals and no tokenizer-attributable format-differential deficit exists in GSM8K's format domain (comma-grouped and bare numbers); partial if digit segmentation is identical but a separator-level tokenizer-attributable deficit exists (H4 confirmed), with GSM8K relevance quantified by the separator counts of S9; confirms only if segmentation differs on bare or comma-grouped numerals and the deficit concentrates accordingly. The GSM8K subsample is sign-only evidence.
- Resources
- Apple silicon with 128 GB unified memory; CPU for the audit; MLX for probes (13 to 17 hours projected, 9.3 hours run) and GSM8K generation (about 1.3 hours run); about 89 GB of model files.
- Prior work
- arXiv:2402.14903 on digit grouping and arithmetic accuracy; arXiv:2402.01035 on digit policies of GPT-2, GPT-4 and Llama; arXiv:2502.19981 on carry errors under single-digit tokenization. arXiv:2604.10799v1 states no digit or punctuation policy for APT4.
Selected exact hypotheses and premises
Premise · P20
85dfefb0-eb84-4deb-b93c-7b494a10b283Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.
Premise · P48
4ab4fa55-c833-4de4-bec3-4d5b68245a5eAPT4's fertility tax on GSM8K questions, below its English preamble tax.