Relative token cost
Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.
Hypothesis
Sign in with GitHubHypothesis · H25 · Author-curated prediction
Tokenizer fertility on three Polish corpora; no model is trained or run.
cited claim · premise
On the Polish text of the Constitution preamble, APT4 has a fertility ratio of 1.62 tokens per word and the Mistral-derived tokenizer 3.22.The Polish fertility gain of a Polish-optimised 32k tokenizer that a science slice could erode.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 0.6461 on Polish PES examination questions (95% interval 0.6420 to 0.6501), above its 0.5020 on the Polish preamble.APT4's fertility tax on one of the three Polish corpora.
Loading research…
Total tokens of the subject tokenizer on the texts divided by the total tokens of the comparator tokenizer on the same texts, minus 1.
The subject's metric against the comparator is below pooled_threshold on the evaluation texts pooled and below per_text_threshold on each of them.
SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.
The Polish language.
SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.
{
"wording": "The science-slice 32k tokenizer's relative token cost against the Polish-only 32k tokenizer is below 0.05 on three Polish corpora pooled and below 0.07 on each of them.",
"predicate": {
"type": "concept",
"key": "cost_below_thresholds"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared with the thresholds.",
"value": {
"type": "concept",
"key": "relative_token_cost"
}
},
{
"role": "subject",
"definition": "Tokenizer whose cost is predicted.",
"value": {
"type": "concept_ref",
"versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
"key": "science_slice_32k_tokenizer"
}
},
{
"role": "comparator",
"definition": "Tokenizer the cost is measured against.",
"value": {
"type": "concept_ref",
"versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
"key": "polish_only_32k_tokenizer"
}
},
{
"role": "evaluation_text",
"definition": "Corpora the cost is measured on.",
"value": {
"type": "text",
"value": "Polish PES examination questions; Polish Wikipedia science articles; Polish reviews"
}
},
{
"role": "pooled_threshold",
"definition": "Upper bound on the pooled cost.",
"value": {
"type": "decimal",
"value": "0.05",
"unit": "ratio"
}
},
{
"role": "per_text_threshold",
"definition": "Upper bound on the cost on each corpus.",
"value": {
"type": "decimal",
"value": "0.07",
"unit": "ratio"
}
},
{
"role": "language",
"definition": "Language of the corpora.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "polish"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →