Finding

Sign in with GitHub
← Publications

Finding · P58 · Author-curated

On numbers of at least four digits grouped with no-break spaces, APT4 uses 10.4214 tokens per number and the Mistral-derived tokenizer 9.0303.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Tokens per numberQuantity compared.
subject
APT4 tokenizerTokenizer after the tokenizer change.
subject value
10.4214 tokens per numberValue of the metric for the subject.
comparator
Mistral-derived tokenizerTokenizer before the tokenizer change.
comparator value
9.0303 tokens per numberValue of the metric for the comparator.
setting
No-break-space-grouped number formatNumber format measured.
scope
synthetic integers from 10^4 to 10^9 and decimals whose integer part has at least four digitsNumbers measured.

Experimental provenance

Method and evaluation protocol
Encode each number after the prefix "a " without special tokens and count the tokens whose character spans overlap it; classify digit policy by comparing token boundaries inside 5,000 integers with single-digit and three-digit chunk predictions.
Dataset
Synthetic number sets S1 to S8 generated with seed 20260703, tokenized by tokenizer files at pinned revisions.Version: unspecified · Access: public
Reported results
10.4214 against 9.0303 tokens per number; APT4 encodes the no-break space as two byte tokens, the Mistral-derived tokenizer as one token.
Uncertainty and replication
Means over deterministic sets.
Limitations
Narrow no-break spaces are three byte tokens under both tokenizers.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Tokens per number

Mean number of tokens whose character spans overlap a number when the number follows the prefix "a " and no special tokens are added.

Key tokens_per_number · version 1235918d-539a-46a7-85fe-2176fb7bcfee

No-break-space-grouped number format

Numbers with thousands grouped by no-break spaces (U+00A0) and a decimal comma.

Key nbsp_number_format · version 1235918d-539a-46a7-85fe-2176fb7bcfee

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Exact references

supports

85dfefb0-eb84-4deb-b93c-7b494a10b283

A special character, the no-break space, changes token efficiency between the two tokenizers.