Finding

Sign in with GitHub
← Publications

Finding · P218 · Author-curated

The relative token cost of the Polish-only 32k tokenizer with separator pieces against the Polish-only 32k tokenizer is -0.000469 on three Polish corpora pooled (95% interval -0.000638 to -0.000325), and 0.000129, -0.001420 and 0.000005 on Polish PES examination questions, Polish Wikipedia science articles and Polish reviews.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric pooled and per corpus

metric
Relative token costQuantity measured.
comparator
Polish-only 32k tokenizerTokenizer the cost is measured against.
evaluation text
Polish PES examination questions, Polish Wikipedia science articles and Polish reviews, pooledCorpora measured.
value
-0.000469 ratioCost on the corpora pooled.
interval low
-0.000638 ratioLower bound of the 95% interval of value.
interval high
-0.000325 ratioUpper bound of the 95% interval of value.
value polish pes questions
0.000129 ratioCost on Polish PES examination questions.
value polish wikipedia science
-0.001420 ratioCost on Polish Wikipedia science articles.
value polish reviews
0.000005 ratioCost on Polish reviews.
language
PolishLanguage of the corpora.

Experimental provenance

Method and evaluation protocol
Sum each tokenizer's tokens over the three Polish corpora, divide by the Polish-only arm's sum and subtract 1; the same per corpus.
Dataset
Three Polish corpora of about 120k words each: PES examination questions, Wikipedia science articles, reviews; two are restricted.Version: unspecified · Access: restricted
Reported results
Pooled -0.000469 [-0.000638, -0.000325]; per corpus PES 0.000129, Wikipedia science -0.001420, reviews 0.000005.
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm. Per-corpus values are points; their intervals are in results/analysis.json.
Limitations
Fertility only; no model is trained on the tokenizers. Each corpus is one sample of its domain. The design, amendments and results were committed together. Polish web text in the training data may overlap the Polish corpora; every arm shares the same Polish text or a prefix of it. The Polish training text skips the first 2,000 documents of the unshuffled FineWeb2-HQ stream, which does not exclude the seed-42 shuffled FineWeb2-HQ Polish holdout: 1 of its 2,000 documents is fully and 43 partly contained in the text; the Polish corpora measured here are not that holdout.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric pooled and per corpus

The metric of the subject against the comparator is value on the evaluation text pooled, with 95% interval interval_low to interval_high, and value_polish_pes_questions, value_polish_wikipedia_science and value_polish_reviews on each of those corpora, in the stated language.

Key metric_pooled_and_per_corpus · version 07fd3382-a184-4c36-9047-709faa73f7d6

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Relative token cost

Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.

Key relative_token_cost · version 2be4b5ff-86ee-4e87-b3a4-77932559f05f

Polish-only 32k tokenizer with separator pieces

SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer on the same Polish text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

Key polish_only_separator_32k_tokenizer · version 6b80001d-30ca-46bf-8aa3-74928a619665

Exact references

supports

6b80001d-30ca-46bf-8aa3-74928a619665

Criterion (b): pooled cost below 0.005.

related

daa6c044-3484-44d0-98d8-e07276dbbd8d

The E5-holdout overlap of the Polish training text.