Finding

Sign in with GitHub
← Publications

Finding · P214 · Author-curated

The relative token cost of the science-slice 32k tokenizer against the Polish-only 32k tokenizer is 0.0215 on three Polish corpora pooled (95% interval 0.0200 to 0.0230), and 0.0258, 0.0180 and 0.0208 on Polish PES examination questions, Polish Wikipedia science articles and Polish reviews.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric pooled and per corpus

metric
Relative token costQuantity measured.
subject
Science-slice 32k tokenizerTokenizer measured.
comparator
Polish-only 32k tokenizerTokenizer the cost is measured against.
evaluation text
Polish PES examination questions, Polish Wikipedia science articles and Polish reviews, pooledCorpora measured.
value
0.0215 ratioCost on the corpora pooled.
interval low
0.0200 ratioLower bound of the 95% interval of value.
interval high
0.0230 ratioUpper bound of the 95% interval of value.
value polish pes questions
0.0258 ratioCost on Polish PES examination questions.
value polish wikipedia science
0.0180 ratioCost on Polish Wikipedia science articles.
value polish reviews
0.0208 ratioCost on Polish reviews.
language
PolishLanguage of the corpora.

Experimental provenance

Method and evaluation protocol
Sum each tokenizer's tokens over the three Polish corpora, divide by the Polish-only arm's sum and subtract 1; the same per corpus.
Dataset
Three Polish corpora of about 120k words each: PES examination questions, Wikipedia science articles, reviews; two are restricted.Version: unspecified · Access: restricted
Reported results
Pooled 0.0215 [0.0200, 0.0230]; per corpus PES 0.0258, Wikipedia science 0.0180, reviews 0.0208.
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm. Per-corpus values are points; their intervals are in results/analysis.json.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together. Polish web text in the training data may overlap the Polish corpora; every arm shares the same Polish text or a prefix of it. The Polish training text skips the first 2,000 documents of the unshuffled FineWeb2-HQ stream, which does not exclude the seed-42 shuffled FineWeb2-HQ Polish holdout: 1 of its 2,000 documents is fully and 43 partly contained in the text; the Polish corpora measured here are not that holdout.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric pooled and per corpus

The metric of the subject against the comparator is value on the evaluation text pooled, with 95% interval interval_low to interval_high, and value_polish_pes_questions, value_polish_wikipedia_science and value_polish_reviews on each of those corpora, in the stated language.

Key metric_pooled_and_per_corpus · version 07fd3382-a184-4c36-9047-709faa73f7d6

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Relative token cost

Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.

Key relative_token_cost · version 2be4b5ff-86ee-4e87-b3a4-77932559f05f

Science-slice 32k tokenizer

SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

Key science_slice_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

supports

2be4b5ff-86ee-4e87-b3a4-77932559f05f

Pooled cost below 0.05 and every per-corpus cost below 0.07.

related

daa6c044-3484-44d0-98d8-e07276dbbd8d

The E5-holdout overlap of the Polish training text.