Finding

Sign in with GitHub
← Publications

Finding · P213 · Author-curated

The pooled English-science excess fertility tax against the Mistral-derived tokenizer is 0.1581 for the balanced 32k tokenizer and 0.5277 for the Polish-only 32k tokenizer, a ratio of 0.2996 (95% interval 0.2823 to 0.3143).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Ratio of metric values

subject
Balanced 32k tokenizerTokenizer measured.
subject value
0.1581 ratioMetric of the subject.
comparator
Polish-only 32k tokenizerTokenizer compared with.
comparator value
0.5277 ratioMetric of the comparator.
ratio
0.2996 ratiosubject_value divided by comparator_value.
interval low
0.2823 ratioLower bound of the 95% interval of ratio.
interval high
0.3143 ratioUpper bound of the 95% interval of ratio.
baseline
Mistral-derived tokenizerTokenizer both metrics are measured against.
evaluation text
English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, pooledCorpora measured.
language
EnglishLanguage of the corpora.

Experimental provenance

Method and evaluation protocol
Sum each tokenizer's tokens over the four English corpora, divide by the Mistral-derived tokenizer's sum and subtract 1; divide the arm's value by the Polish-only arm's.
Dataset
Four English corpora of about 120k words each: arXiv abstracts, LaTeX method sections, Python code, GSM8K problems; two are restricted.Version: unspecified · Access: restricted
Reported results
R 0.2996 [0.2823, 0.3143]; one-sided bootstrap p(R >= 1) 0.000999; pooled excess tax 0.1581 [0.1458, 0.1692] against 0.5277 [0.5131, 0.5410].
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Balanced 32k tokenizer

SentencePiece BPE tokenizer built like the science-slice 32k tokenizer but trained on 60% of the same Polish web text, 20% general English web text from SlimPajama-6B and 20% of the same science text.

Key balanced_32k_tokenizer · version 2e436203-2e44-4b97-8690-d61753e660fb

Ratio of metric values

subject_value, the metric of the subject, divided by comparator_value, the metric of the comparator, is ratio, with 95% paired-bootstrap interval interval_low to interval_high; both metrics are measured against the baseline tokenizer on the evaluation text in the stated language.

Key ratio_of_metric · version c7cdbb52-799a-4c5b-a448-9597a83d0ccd

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Pooled English-science excess fertility tax

Total tokens of a tokenizer over four English corpora (arXiv abstracts, LaTeX method sections, Python code, GSM8K problems) divided by the total tokens of the baseline tokenizer over the same corpora, minus 1.

Key pooled_excess_science_tax · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

related

d1a46254-542b-43a9-b64b-3536a37e2efd

Exploratory arm without a bin.