Finding

Sign in with GitHub
← Publications

Finding · P223 · Author-curated

On GSM8K problems, the excess fertility tax against the Mistral-derived tokenizer of the science-slice 32k tokenizer is 0.4713 times that of the Polish-only 32k tokenizer (95% interval 0.4667 to 0.4757).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Ratio with interval

metric
Excess fertility taxQuantity compared.
subject
Science-slice 32k tokenizerTokenizer measured.
comparator
Polish-only 32k tokenizerTokenizer compared with.
ratio
0.4713 ratioSubject's metric divided by the comparator's.
interval low
0.4667 ratioLower bound of the 95% interval.
interval high
0.4757 ratioUpper bound of the 95% interval.
baseline
Mistral-derived tokenizerTokenizer both metrics are measured against.
evaluation text
GSM8K problemsCorpus measured.
language
EnglishLanguage of the corpus.

Experimental provenance

Method and evaluation protocol
Divide the arm's tokens on the corpus by the Mistral-derived tokenizer's and subtract 1; divide by the same value for the Polish-only arm.
Dataset
GSM8K problems.Version: unspecified · Access: public
Reported results
R 0.4713 [0.4667, 0.4757]; Holm-adjusted one-sided bootstrap p(R >= 1) at most 0.003996 across the four English corpora.
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm. The interval is not adjusted for the four corpora; the p-value is.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Ratio with interval

ratio is the subject's metric divided by the comparator's metric, with 95% paired-bootstrap interval interval_low to interval_high, both measured against the baseline tokenizer on the evaluation text in the stated language.

Key ratio_with_interval · version 55822f7d-0815-4185-b506-002e4d1767fd

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

GSM8K problems

The 7,473 questions of the GSM8K training split.

Key english_gsm8k_problems · version 4ab4fa55-c833-4de4-bec3-4d5b68245a5e

Excess fertility tax

Tokens of a tokenizer on a text divided by tokens of the baseline tokenizer on the same text, minus 1.

Key excess_fertility_tax · version 55822f7d-0815-4185-b506-002e4d1767fd

Science-slice 32k tokenizer

SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

Key science_slice_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

extends

4ab4fa55-c833-4de4-bec3-4d5b68245a5e

APT4's tax on the same corpus.