Finding

Sign in with GitHub
← Publications

Finding · P215 · Author-curated

The relative token cost of the science-slice 32k tokenizer with whitespace pieces against the Polish-only 32k tokenizer is 0.0179 on three Polish corpora pooled (95% interval 0.0116 to 0.0222), and 0.0258, 0.0080 and 0.0207 on Polish PES examination questions, Polish Wikipedia science articles and Polish reviews.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric pooled and per corpus

metric
Relative token costQuantity measured.
comparator
Polish-only 32k tokenizerTokenizer the cost is measured against.
evaluation text
Polish PES examination questions, Polish Wikipedia science articles and Polish reviews, pooledCorpora measured.
value
0.0179 ratioCost on the corpora pooled.
interval low
0.0116 ratioLower bound of the 95% interval of value.
interval high
0.0222 ratioUpper bound of the 95% interval of value.
value polish pes questions
0.0258 ratioCost on Polish PES examination questions.
value polish wikipedia science
0.0080 ratioCost on Polish Wikipedia science articles.
value polish reviews
0.0207 ratioCost on Polish reviews.
language
PolishLanguage of the corpora.

Experimental provenance

Method and evaluation protocol
Sum each tokenizer's tokens over the three Polish corpora, divide by the Polish-only arm's sum and subtract 1; the same per corpus.
Dataset
Three Polish corpora of about 120k words each: PES examination questions, Wikipedia science articles, reviews; two are restricted.Version: unspecified · Access: restricted
Reported results
Pooled 0.0179 [0.0116, 0.0222]; per corpus PES 0.0258, Wikipedia science 0.0080, reviews 0.0207.
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm. Per-corpus values are points; their intervals are in results/analysis.json.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together. Polish web text in the training data may overlap the Polish corpora; every arm shares the same Polish text or a prefix of it. The Polish training text skips the first 2,000 documents of the unshuffled FineWeb2-HQ stream, which does not exclude the seed-42 shuffled FineWeb2-HQ Polish holdout: 1 of its 2,000 documents is fully and 43 partly contained in the text; the Polish corpora measured here are not that holdout.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric pooled and per corpus

The metric of the subject against the comparator is value on the evaluation text pooled, with 95% interval interval_low to interval_high, and value_polish_pes_questions, value_polish_wikipedia_science and value_polish_reviews on each of those corpora, in the stated language.

Key metric_pooled_and_per_corpus · version 07fd3382-a184-4c36-9047-709faa73f7d6

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Relative token cost

Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.

Key relative_token_cost · version 2be4b5ff-86ee-4e87-b3a4-77932559f05f

Science-slice 32k tokenizer with whitespace pieces

SentencePiece BPE tokenizer built like the science-slice 32k tokenizer on the same text, with whitespace-only pieces allowed.

Key science_slice_whitespace_32k_tokenizer · version 59656f61-c3e1-480c-91d0-4e1f202980cf

Exact references

related

2be4b5ff-86ee-4e87-b3a4-77932559f05f

The whitespace arm under the same thresholds.

related

daa6c044-3484-44d0-98d8-e07276dbbd8d

The E5-holdout overlap of the Polish training text.