At most a fraction of
The subject's metric against the baseline tokenizer is at most fraction times the comparator's metric against the same baseline.
Hypothesis
Sign in with GitHubHypothesis · H24 · Author-curated prediction
Tokenizer fertility of 32k tokenizers trained on 1 GiB of text each, measured on four English corpora; no model is trained or run.
finding · premise
APT4's English fertility tax exceeds its preamble tax on 2 of 4 scientific corpora (English arXiv abstracts and English Python code) and is below it on the other 2 (English LaTeX method sections and GSM8K problems).APT4's English fertility tax depends on the science domain.
cited claim · premise
APT4's vocabulary size is kept at approximately 32k tokens to isolate improvements from segmentation efficiency rather than from increased vocabulary capacity.APT4's vocabulary is held at about 32k tokens, so allocation within that size is the lever.
Loading research…
The subject's metric against the baseline tokenizer is at most fraction times the comparator's metric against the same baseline.
SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.
Total tokens of a tokenizer over four English corpora (arXiv abstracts, LaTeX method sections, Python code, GSM8K problems) divided by the total tokens of the baseline tokenizer over the same corpora, minus 1.
SentencePiece BPE tokenizer built like the Polish-only 32k tokenizer but trained on 1 GiB of text that is 80% the same Polish web text and 20% science text in equal parts arXiv LaTeX, Python code and mathematical web text, with the no-break space, narrow no-break space and minus sign as fixed single-character pieces.
Tokenizer of the original Bielik v3 models, derived from Mistral's.
The English language.
{
"wording": "The science-slice 32k tokenizer's pooled English-science excess fertility tax against the Mistral-derived tokenizer is at most half that of the Polish-only 32k tokenizer.",
"predicate": {
"type": "concept",
"key": "at_most_fraction_of"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "pooled_excess_science_tax"
}
},
{
"role": "subject",
"definition": "Tokenizer predicted to have the lower value.",
"value": {
"type": "concept",
"key": "science_slice_32k_tokenizer"
}
},
{
"role": "comparator",
"definition": "Tokenizer compared with.",
"value": {
"type": "concept",
"key": "polish_only_32k_tokenizer"
}
},
{
"role": "baseline",
"definition": "Tokenizer both metrics are measured against.",
"value": {
"type": "concept_ref",
"versionId": "bd3d8329-e626-4805-9fff-f83f146769a7",
"key": "mistral_tokenizer"
}
},
{
"role": "fraction",
"definition": "Largest predicted ratio of the subject's value to the comparator's.",
"value": {
"type": "decimal",
"value": "0.50",
"unit": "ratio"
}
},
{
"role": "language",
"definition": "Language of the evaluation texts.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "english"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →