One token at a cost below a threshold
Each of the characters encodes as one token between digits with the subject tokenizer, and the subject's metric against the comparator on the evaluation texts pooled is below threshold.
Hypothesis
Sign in with GitHubHypothesis · H26 · Author-curated prediction
Digit and separator tokenization on synthetic number renderings and tokenizer fertility on three Polish corpora; no model is trained or run.
cited claim · premise
Beyond vocabulary size, the handling of digits, punctuation and special characters can influence both token efficiency and downstream generation quality.Handling of digits, punctuation and special characters can change token efficiency.
finding · premise
On numbers of at least four digits grouped with no-break spaces, APT4 uses 10.4214 tokens per number and the Mistral-derived tokenizer 9.0303.APT4 needs more tokens than the Mistral-derived tokenizer on numbers grouped with no-break spaces.
finding · premise
Bielik-PL-11B-v3.0-Instruct's arithmetic accuracy minus Bielik-11B-v3.0-Instruct's with space-grouped numbers, minus the same difference with no-break-space-grouped numbers, is 0.1736 (95% interval 0.1535 to 0.1956; Holm-adjusted bootstrap p-value 0 with 1,000 resamples).The model with APT4 loses arithmetic accuracy on no-break-space-grouped numbers relative to space-grouped ones, beyond the original model.
Loading research…
Each of the characters encodes as one token between digits with the subject tokenizer, and the subject's metric against the comparator on the evaluation texts pooled is below threshold.
SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer on the same Polish text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.
SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.
The Polish language.
Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.
{
"wording": "With the no-break space, narrow no-break space and minus sign added as fixed pieces, the Polish-only 32k tokenizer encodes each of them as one token between digits at a pooled relative token cost below 0.005 on three Polish corpora.",
"predicate": {
"type": "concept",
"key": "one_token_at_cost_below"
},
"roles": [
{
"role": "subject",
"definition": "Tokenizer with the fixed pieces.",
"value": {
"type": "concept",
"key": "polish_only_separator_32k_tokenizer"
}
},
{
"role": "comparator",
"definition": "Tokenizer without them, trained on the same text.",
"value": {
"type": "concept_ref",
"versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
"key": "polish_only_32k_tokenizer"
}
},
{
"role": "characters",
"definition": "Characters given fixed pieces.",
"value": {
"type": "text",
"value": "U+00A0 no-break space; U+202F narrow no-break space; U+2212 minus sign"
}
},
{
"role": "metric",
"definition": "Cost compared with the threshold.",
"value": {
"type": "concept_ref",
"versionId": "2be4b5ff-86ee-4e87-b3a4-77932559f05f",
"key": "relative_token_cost"
}
},
{
"role": "threshold",
"definition": "Upper bound on the pooled cost.",
"value": {
"type": "decimal",
"value": "0.005",
"unit": "ratio"
}
},
{
"role": "evaluation_text",
"definition": "Corpora the cost is measured on.",
"value": {
"type": "text",
"value": "Polish PES examination questions; Polish Wikipedia science articles; Polish reviews"
}
},
{
"role": "language",
"definition": "Language of the corpora.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "polish"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →