Finding

Sign in with GitHub
← Publications

Finding · P56 · Author-curated

Of the 9 tokenizers of Table 1 of arXiv:2604.10799v1, 7 split every classified integer into one token per digit, the SmolLM3 (Llama 3.2) tokenizer splits every one into three-digit chunks from the left, and APT3 follows neither pattern.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Digit policy class counts

tokenizers total
9 tokenizersTokenizers classified.
count single digit
7 tokenizersTokenizers with the single-digit policy.
single digit tokenizers
APT4; Mistral-derived tokenizer; Qwen3; Gemma 3; EuroLLM; Mistral Small 3.1; ApertusTokenizers with the single-digit policy.
count three digit left to right
1 tokenizersTokenizers splitting digits into three-digit chunks from the left.
three digit left to right tokenizers
SmolLM3 (Llama 3.2)Tokenizers splitting digits into three-digit chunks from the left.
count free bpe
1 tokenizersTokenizers following neither pattern.
free bpe tokenizers
APT3Tokenizers following neither pattern.
integers classified
5000 integersIntegers classified per tokenizer.

Experimental provenance

Method and evaluation protocol
Encode each number after the prefix "a " without special tokens and count the tokens whose character spans overlap it; classify digit policy by comparing token boundaries inside 5,000 integers with single-digit and three-digit chunk predictions.
Dataset
Synthetic number sets S1 to S8 generated with seed 20260703, tokenized by tokenizer files at pinned revisions.Version: unspecified · Access: public
Reported results
Single-digit: APT4; Mistral-derived tokenizer; Qwen3; Gemma 3; EuroLLM; Mistral Small 3.1; Apertus. Three-digit chunks from the left: SmolLM3 (Llama 3.2). Neither: APT3. Gemma 3 and EuroLLM have multi-digit vocabulary tokens (176 and 4) but split every classified integer one token per digit.
Uncertainty and replication
Exact counts over deterministic sets.
Limitations
The Apertus and Mistral Small 3.1 tokenizers have identical audit profiles although their tokenizer files differ.

Author’s note

Class counts computed by scripts/09_metrics.py from results/audit.json.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Digit policy class counts

Of tokenizers_total tokenizers, count_single_digit (single_digit_tokenizers) split every one of integers_classified integers one token per digit, count_three_digit_left_to_right (three_digit_left_to_right_tokenizers) split every one into three-digit chunks from the left, and count_free_bpe (free_bpe_tokenizers) follow neither pattern.

Key digit_policy_class_counts · version 93a0806d-e3fc-4d91-b07b-c04c13cb793f