Finding

Sign in with GitHub
← Publications

Finding · P222 · Author-curated

On English Python code, the excess fertility tax against the Mistral-derived tokenizer of the science-slice 32k tokenizer is 0.5763 times that of the Polish-only 32k tokenizer (95% interval 0.5594 to 0.5939).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Ratio with interval

metric
Excess fertility taxQuantity compared.
subject
Science-slice 32k tokenizerTokenizer measured.
comparator
Polish-only 32k tokenizerTokenizer compared with.
ratio
0.5763 ratioSubject's metric divided by the comparator's.
interval low
0.5594 ratioLower bound of the 95% interval.
interval high
0.5939 ratioUpper bound of the 95% interval.
baseline
Mistral-derived tokenizerTokenizer both metrics are measured against.
evaluation text
English Python codeCorpus measured.
language
EnglishLanguage of the corpus.

Experimental provenance

Method and evaluation protocol
Divide the arm's tokens on the corpus by the Mistral-derived tokenizer's and subtract 1; divide by the same value for the Polish-only arm.
Dataset
English Python code.Version: unspecified · Access: restricted
Reported results
R 0.5763 [0.5594, 0.5939]; Holm-adjusted one-sided bootstrap p(R >= 1) at most 0.003996 across the four English corpora.
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm. The interval is not adjusted for the four corpora; the p-value is.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Ratio with interval

ratio is the subject's metric divided by the comparator's metric, with 95% paired-bootstrap interval interval_low to interval_high, both measured against the baseline tokenizer on the evaluation text in the stated language.

Key ratio_with_interval · version 55822f7d-0815-4185-b506-002e4d1767fd

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

English Python code

331 Python files from GitHub, truncated at 700 words.

Key english_python_code · version 4acc878a-a4a3-45bd-bb03-2cd8541a7856

Excess fertility tax

Tokens of a tokenizer on a text divided by tokens of the baseline tokenizer on the same text, minus 1.

Key excess_fertility_tax · version 55822f7d-0815-4185-b506-002e4d1767fd

Science-slice 32k tokenizer

SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

Key science_slice_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

extends

4acc878a-a4a3-45bd-bb03-2cd8541a7856

APT4's tax on the same corpus.