Finding

Sign in with GitHub
← Publications

Finding · P211 · Author-curated

The pooled English-science excess fertility tax against the Mistral-derived tokenizer is 0.2109 for the science-slice 32k tokenizer and 0.5277 for the Polish-only 32k tokenizer, a ratio of 0.3997 (95% interval 0.3838 to 0.4135).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Ratio of metric values

subject
Science-slice 32k tokenizerTokenizer measured.
subject value
0.2109 ratioMetric of the subject.
comparator
Polish-only 32k tokenizerTokenizer compared with.
comparator value
0.5277 ratioMetric of the comparator.
ratio
0.3997 ratiosubject_value divided by comparator_value.
interval low
0.3838 ratioLower bound of the 95% interval of ratio.
interval high
0.4135 ratioUpper bound of the 95% interval of ratio.
baseline
Mistral-derived tokenizerTokenizer both metrics are measured against.
evaluation text
English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, pooledCorpora measured.
language
EnglishLanguage of the corpora.

Experimental provenance

Method and evaluation protocol
Sum each tokenizer's tokens over the four English corpora, divide by the Mistral-derived tokenizer's sum and subtract 1; divide the arm's value by the Polish-only arm's.
Dataset
Four English corpora of about 120k words each: arXiv abstracts, LaTeX method sections, Python code, GSM8K problems; two are restricted.Version: unspecified · Access: restricted
Reported results
R 0.3997 [0.3838, 0.4135]; one-sided bootstrap p(R >= 1) 0.000999; pooled excess tax 0.2109 [0.1980, 0.2225] against 0.5277 [0.5131, 0.5410].
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Ratio of metric values

subject_value, the metric of the subject, divided by comparator_value, the metric of the comparator, is ratio, with 95% paired-bootstrap interval interval_low to interval_high; both metrics are measured against the baseline tokenizer on the evaluation text in the stated language.

Key ratio_of_metric · version c7cdbb52-799a-4c5b-a448-9597a83d0ccd

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Pooled English-science excess fertility tax

Total tokens of a tokenizer over four English corpora (arXiv abstracts, LaTeX method sections, Python code, GSM8K problems) divided by the total tokens of the baseline tokenizer over the same corpora, minus 1.

Key pooled_excess_science_tax · version d1a46254-542b-43a9-b64b-3536a37e2efd

Science-slice 32k tokenizer

SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

Key science_slice_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

supports

d1a46254-542b-43a9-b64b-3536a37e2efd

The ratio is at most 0.50.

extends

3d1e1910-f471-4b1a-be65-5e2db901f9a0

Measures how training-data allocation changes the English-science tax at the same vocabulary size.