Finding

Sign in with GitHub
← Publications

Finding · P225 · Author-curated

On English LaTeX method sections, the excess fertility tax against the Mistral-derived tokenizer of the science-slice 32k tokenizer with whitespace pieces is 0.0674 times that of the Polish-only 32k tokenizer (95% interval 0.0595 to 0.0760).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Ratio with interval

metric
Excess fertility taxQuantity compared.
comparator
Polish-only 32k tokenizerTokenizer compared with.
ratio
0.0674 ratioSubject's metric divided by the comparator's.
interval low
0.0595 ratioLower bound of the 95% interval.
interval high
0.0760 ratioUpper bound of the 95% interval.
baseline
Mistral-derived tokenizerTokenizer both metrics are measured against.
evaluation text
English LaTeX method sectionsCorpus measured.
language
EnglishLanguage of the corpus.

Experimental provenance

Method and evaluation protocol
Divide the arm's tokens on the corpus by the Mistral-derived tokenizer's and subtract 1; divide by the same value for the Polish-only arm.
Dataset
English LaTeX method sections.Version: unspecified · Access: restricted
Reported results
R 0.0674 [0.0595, 0.0760]; Holm-adjusted one-sided bootstrap p(R >= 1) at most 0.003996 across the four English corpora.
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm. The interval is not adjusted for the four corpora; the p-value is.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Ratio with interval

ratio is the subject's metric divided by the comparator's metric, with 95% paired-bootstrap interval interval_low to interval_high, both measured against the baseline tokenizer on the evaluation text in the stated language.

Key ratio_with_interval · version 55822f7d-0815-4185-b506-002e4d1767fd

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

English LaTeX method sections

219 LaTeX sections of arXiv papers whose titles name methods, 80 to 1500 words per paper.

Key english_latex_methods · version 683761d9-7e85-4d6e-9a95-bd924d442e75

Excess fertility tax

Tokens of a tokenizer on a text divided by tokens of the baseline tokenizer on the same text, minus 1.

Key excess_fertility_tax · version 55822f7d-0815-4185-b506-002e4d1767fd

Science-slice 32k tokenizer with whitespace pieces

SentencePiece BPE tokenizer built like the science-slice 32k tokenizer on the same text, with whitespace-only pieces allowed.

Key science_slice_whitespace_32k_tokenizer · version 59656f61-c3e1-480c-91d0-4e1f202980cf

Exact references

extends

683761d9-7e85-4d6e-9a95-bd924d442e75

APT4's tax on the same corpus.