Finding

Sign in with GitHub
← Publications

Finding · P219 · Author-curated

At a fixed 32,000-token vocabulary, the science-slice 32k tokenizer has at most half the pooled English-science excess fertility tax of the Polish-only 32k tokenizer (ratio 0.3997) at a relative Polish token cost below 0.05 pooled (0.0215) and below 0.07 on each Polish corpus (at most 0.0258).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Meets a decision rule

subject
Science-slice 32k tokenizerTokenizer the rule is applied to.
comparator
Polish-only 32k tokenizerTokenizer compared with.
first result
Exact publication c7cdbb52-799a-4c5b-a448-9597a83d0ccdResult on the English-science tax.
second result
Exact publication 07fd3382-a184-4c36-9047-709faa73f7d6Result on the Polish token cost.
rule
ratio of pooled English-science excess fertility taxes at most 0.50; pooled Polish relative token cost below 0.05; each per-corpus Polish relative token cost below 0.07Thresholds applied.

Experimental provenance

Method and evaluation protocol
Apply the locked decision rule to the pooled ratio and to the pooled and per-corpus Polish costs of the science-slice arm.
Dataset
Seven English and Polish corpora of about 120k words each; four are restricted.Version: unspecified · Access: restricted
Reported results
H1 bin confirmed (R 0.3997); H2 bin confirmed (pooled 0.0215, largest per corpus 0.0258); verdict confirmed. The whitespace arm meets the same thresholds.
Uncertainty and replication
Bins read on point estimates; every relevant 95% interval lies on the same side of its threshold.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together. The rule was committed with the results.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Meets a decision rule

The results in first_result and second_result satisfy every threshold in rule for the subject against the comparator.

Key meets_rule · version fd49148b-f337-4f5a-a26d-4c8227d57f1e

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Science-slice 32k tokenizer

SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

Key science_slice_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

derived from

c7cdbb52-799a-4c5b-a448-9597a83d0ccd

The tax ratio the rule reads.

derived from

07fd3382-a184-4c36-9047-709faa73f7d6

The Polish costs the rule reads.

related

e6daa427-c8d9-4276-bc5b-aacb1fc82344

Fertility changes with training-data allocation at a fixed vocabulary size.