Finding

Sign in with GitHub
← Publications

Finding · P227 · Author-curated

On numbers of at least four digits with no-break-space digit-group separators, the Polish-only 32k tokenizer needs 10.4214 tokens per number and APT4 10.4214.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Tokens per numberQuantity compared.
subject
Polish-only 32k tokenizerTokenizer measured.
subject value
10.4214 tokens per numberValue of the metric for the subject.
comparator
APT4 tokenizerTokenizer compared with.
comparator value
10.4214 tokens per numberValue of the metric for the comparator.
scope
synthetic integers from 10^4 to 10^9 and decimals whose integer part has at least four digitsNumbers measured.

Experimental provenance

Method and evaluation protocol
Mean tokens over the number span for every rendering with no-break-space separators and at least four digits in the rendering audit set.
Dataset
Synthetic number renderings; no corpus.Version: unspecified · Access: public
Reported results
Polish-only arm 10.4214; APT4 10.4214; separator arm 9.0303; Mistral-derived 9.0303.
Uncertainty and replication
Exact means over a fixed synthetic set; no sampling.
Limitations
Synthetic renderings only; frequency of such numbers in real text is not measured here.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Tokens per number

Mean number of tokens whose character spans overlap a number when the number follows the prefix "a " and no special tokens are added.

Key tokens_per_number · version 1235918d-539a-46a7-85fe-2176fb7bcfee

No-break-space-grouped number format

Numbers with thousands grouped by no-break spaces (U+00A0) and a decimal comma.

Key nbsp_number_format · version 1235918d-539a-46a7-85fe-2176fb7bcfee

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

extends

3d25bdbe-b09d-4c1a-9e65-5fa802f51d9d

The same measurement on the same numbers gives APT4 10.4214; a Polish-only tokenizer trained from scratch reaches the same value.

related

85dfefb0-eb84-4deb-b93c-7b494a10b283

Token cost of a special character in numbers.