Finding

Sign in with GitHub
← Publications

Finding · P217 · Author-curated

Between digits, the no-break space, narrow no-break space and minus sign add 1, 1 and 1 tokens with the Polish-only 32k tokenizer with separator pieces, 2, 3 and 3 with the Polish-only 32k tokenizer, 2, 3 and 3 with APT4, and 1, 3 and 1 with the Mistral-derived tokenizer.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Tokens per character by tokenizer

metric
Tokens added between digitsQuantity measured.
characters
U+00A0 no-break space; U+202F narrow no-break space; U+2212 minus signCharacters measured, in the order of every values role.
subject
Polish-only 32k tokenizer with separator piecesTokenizer with fixed pieces for the characters.
subject values
1; 1; 1Tokens per character with the subject.
comparator
Polish-only 32k tokenizerSame-text tokenizer without the pieces.
comparator values
2; 3; 3Tokens per character with the comparator.
reference
APT4 tokenizerFirst reference tokenizer.
reference values
2; 3; 3Tokens per character with the first reference.
second reference
Mistral-derived tokenizerSecond reference tokenizer.
second reference values
1; 3; 1Tokens per character with the second reference.

Experimental provenance

Method and evaluation protocol
Encode '1<character>234' and '1234' after the carrier 'a ' without special tokens and count the tokens that intersect the number span.
Dataset
Synthetic number strings; no corpus.Version: unspecified · Access: public
Reported results
Separator arm 1; 1; 1; Polish-only arm 2; 3; 3; APT4 2; 3; 3; Mistral-derived 1; 3; 1. The Polish-only arm learned none of the three characters as pieces.
Uncertainty and replication
Exact counts; no sampling.
Limitations
Fixed pieces act as segmentation boundaries during training, so no merge includes them (no fused word-boundary minus piece). The audit script was last written seconds before its only recorded output.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Tokens added between digits

Tokens in the encoding of '1', the character and '234' minus tokens in the encoding of '1234', each after the carrier text 'a ' and without special tokens.

Key tokens_added_between_digits · version 8a4b7bae-a4a9-406e-ab8f-3ef8d0766aaf

Tokens per character by tokenizer

For each character in characters, in order, the metric is the matching entry of subject_values for the subject, comparator_values for the comparator, reference_values for the reference and second_reference_values for the second reference.

Key tokens_per_character_by_tokenizer · version 8a4b7bae-a4a9-406e-ab8f-3ef8d0766aaf

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Polish-only 32k tokenizer with separator pieces

SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer on the same Polish text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.

Key polish_only_separator_32k_tokenizer · version 6b80001d-30ca-46bf-8aa3-74928a619665

Exact references

supports

6b80001d-30ca-46bf-8aa3-74928a619665

Criterion (a): each character is one token between digits.

extends

85dfefb0-eb84-4deb-b93c-7b494a10b283

Measures the token cost of three special characters in numbers with and without fixed pieces.

extends

bd8acd13-639f-4b0d-bb2b-7a180e1f837c

APT4 and the Mistral-derived tokenizer segment no-break-space and U+2212 exemplars differently; fixed pieces make each character one token.

related

81582930-04ae-4afe-809e-f2d13cec4c18

The arithmetic accuracy cost of no-break-space grouping for the model with APT4, which the pieces target.