Finding

Sign in with GitHub
← Publications

Finding · P212 · Author-curated

The pooled English-science excess fertility tax against the Mistral-derived tokenizer is 0.0884 for the science-slice 32k tokenizer with whitespace pieces and 0.5277 for the Polish-only 32k tokenizer, a ratio of 0.1675 (95% interval 0.1617 to 0.1741).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Ratio of metric values

subject value
0.0884 ratioMetric of the subject.
comparator
Polish-only 32k tokenizerTokenizer compared with.
comparator value
0.5277 ratioMetric of the comparator.
ratio
0.1675 ratiosubject_value divided by comparator_value.
interval low
0.1617 ratioLower bound of the 95% interval of ratio.
interval high
0.1741 ratioUpper bound of the 95% interval of ratio.
baseline
Mistral-derived tokenizerTokenizer both metrics are measured against.
evaluation text
English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, pooledCorpora measured.
language
EnglishLanguage of the corpora.

Experimental provenance

Method and evaluation protocol
Sum each tokenizer's tokens over the four English corpora, divide by the Mistral-derived tokenizer's sum and subtract 1; divide the arm's value by the Polish-only arm's.
Dataset
Four English corpora of about 120k words each: arXiv abstracts, LaTeX method sections, Python code, GSM8K problems; two are restricted.Version: unspecified · Access: restricted
Reported results
R 0.1675 [0.1617, 0.1741]; one-sided bootstrap p(R >= 1) 0.000999; pooled excess tax 0.0884 [0.0852, 0.0914] against 0.5277 [0.5131, 0.5410].
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Science-slice 32k tokenizer with whitespace pieces

SentencePiece BPE tokenizer built like the science-slice 32k tokenizer on the same text, with whitespace-only pieces allowed.

Key science_slice_whitespace_32k_tokenizer · version 59656f61-c3e1-480c-91d0-4e1f202980cf

Ratio of metric values

subject_value, the metric of the subject, divided by comparator_value, the metric of the comparator, is ratio, with 95% paired-bootstrap interval interval_low to interval_high; both metrics are measured against the baseline tokenizer on the evaluation text in the stated language.

Key ratio_of_metric · version c7cdbb52-799a-4c5b-a448-9597a83d0ccd

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Pooled English-science excess fertility tax

Total tokens of a tokenizer over four English corpora (arXiv abstracts, LaTeX method sections, Python code, GSM8K problems) divided by the total tokens of the baseline tokenizer over the same corpora, minus 1.

Key pooled_excess_science_tax · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

related

d1a46254-542b-43a9-b64b-3536a37e2efd

Secondary arm under the same bins.