Finding

Sign in with GitHub
← Publications

Finding · P54 · Author-curated

APT4 splits all 5,000 classified integers into one token per digit and has no multi-digit tokens in its vocabulary.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Digit policy measurement

subject
APT4 tokenizerTokenizer measured.
policy
Single-digit policyDigit policy class.
share single digit
1 share of integersShare of classified integers split one token per digit.
share three digit left to right
0 share of integersShare of classified integers split into three-digit chunks from the left.
integers classified
5000 integersIntegers classified.
multi digit tokens
0 tokensVocabulary tokens of two or more digits.
max digit run
1 digitsLongest digit run in one vocabulary token.
place value boundary coverage
1 share of integersShare of integers whose thousands-group boundaries are all token boundaries.
single token coverage 0 999
0.01 share of integersShare of the integers 0 to 999 encoded as one token.

Experimental provenance

Method and evaluation protocol
Encode each number after the prefix "a " without special tokens and count the tokens whose character spans overlap it; classify digit policy by comparing token boundaries inside 5,000 integers with single-digit and three-digit chunk predictions.
Dataset
Synthetic number sets S1 to S8 generated with seed 20260703, tokenized by tokenizer files at pinned revisions.Version: unspecified · Access: public
Reported results
Single-digit share 1; multi-digit vocabulary tokens 0; longest digit token 1; place-value boundary coverage 1.
Uncertainty and replication
Exact counts over deterministic sets.
Limitations
The classified integers are 3,000 of 2 to 4 digits and 2,000 of 5 to 10 digits; the hypothesis this supports was written after the tokenizer files had been inspected.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Digit policy measurement

The subject tokenizer's digit segmentation belongs to the policy class. Of integers_classified integers, share_single_digit are split one token per digit and share_three_digit_left_to_right into three-digit chunks from the left; multi_digit_tokens counts vocabulary tokens of two or more digits and max_digit_run is the longest digit run in one token; place_value_boundary_coverage is the share of integers whose thousands-group boundaries are all token boundaries; single_token_coverage_0_999 is the share of the integers 0 to 999 encoded as one token.

Key digit_policy_measurement · version 2f04e3ac-4c60-437e-93d3-0e5644c55b44

Single-digit policy

Segmentation of every digit of a digit string into its own token.

Key single_digit_policy · version a72eb751-c3d6-4335-ba12-2776c48ff646

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Exact references

related

85dfefb0-eb84-4deb-b93c-7b494a10b283

States APT4's digit segmentation, which the paper does not specify.