Residual ratio
The subject arm's bits per byte minus the reference arm's bits per byte, pooled over the texts, divided by the comparator arm's bits per byte minus the reference arm's bits per byte on the same texts.
Hypothesis
Sign in with GitHubHypothesis · H48 · Author-curated prediction
Five arms trained on one deterministic document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B text by Qwen tokens, the sequence of the Qwen2.5-1.5B continued-pretraining arms: the two named arms, the Polish-only 32k tokenizer with separator pieces, the science-slice 32k tokenizer and the Qwen-tokenizer reference; one run per arm; byte-anchored bits per byte; matched-step and matched-byte comparisons both reported. A ratio of at least 0.9 means the allocation buys nothing in loss; at most 0.2 means full transfer.
finding · premise
The pooled English-science excess fertility tax against the Mistral-derived tokenizer is 0.0884 for the science-slice 32k tokenizer with whitespace pieces and 0.5277 for the Polish-only 32k tokenizer, a ratio of 0.1675 (95% interval 0.1617 to 0.1741).The fertility-tax ratio of the same two tokenizers.
finding · premise
The relative token cost of the science-slice 32k tokenizer with whitespace pieces against the Polish-only 32k tokenizer is 0.0179 on three Polish corpora pooled (95% interval 0.0116 to 0.0222), and 0.0258, 0.0080 and 0.0207 on Polish PES examination questions, Polish Wikipedia science articles and Polish reviews.The Polish token cost of the whitespace science-slice tokenizer.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on Python code is 0.5941 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.3795 for Qwen2.5-1.5B with its own tokenizer.A formal-text residual of the APT4 arm at matched continued pretraining.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on math_clean arithmetic statements is 1.7672 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 1.5157 for Qwen2.5-1.5B with its own tokenizer.A formal-text residual of the APT4 arm at matched continued pretraining.
Loading research…
The subject arm's bits per byte minus the reference arm's bits per byte, pooled over the texts, divided by the comparator arm's bits per byte minus the reference arm's bits per byte on the same texts.
The metric of subject against comparator on the texts is at most threshold, and the cost_metric of subject exceeds that of comparator by at most cost_limit, where subject and comparator are the tokenizers given to the base_model by the initialisation before the procedure for the budget, and reference_arm is the base_model with its own tokenizer after the same procedure.
The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.
SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.
Bits per byte over the pooled documents of four Polish texts, the first 200 documents of each: a FineWeb2-HQ holdout, Polish Wikipedia science articles, Polish PES examination questions and Polish reviews.
An embedding initialisation for a replaced tokenizer's vocabulary.
Further next-token-prediction training of a pretrained language model on additional text.
SentencePiece BPE tokenizer built like the science-slice 32k tokenizer on the same text, with whitespace-only pieces allowed.
English LaTeX method sections, English Python code and math_clean statements.
{
"wording": "At matched continued pretraining of 1B tokens, transplanting the science-slice 32k tokenizer with whitespace pieces into Qwen2.5-0.5B by Fast Vocabulary Transfer, instead of the Polish-only 32k tokenizer, leaves at most half of the pooled bits-per-byte residual on English LaTeX method sections, Python code and math_clean statements over Qwen2.5-0.5B continued with its own tokenizer, at a pooled Polish bits-per-byte cost of at most 0.02.",
"predicate": {
"type": "concept",
"key": "residual_ratio_and_cost_at_most"
},
"roles": [
{
"role": "subject",
"definition": "Tokenizer of the arm predicted to keep less residual.",
"value": {
"type": "concept_ref",
"versionId": "59656f61-c3e1-480c-91d0-4e1f202980cf",
"key": "science_slice_whitespace_32k_tokenizer"
}
},
{
"role": "comparator",
"definition": "Tokenizer of the arm it is compared with.",
"value": {
"type": "concept_ref",
"versionId": "d1a46254-542b-43a9-b64b-3536a37e2efd",
"key": "polish_only_32k_tokenizer"
}
},
{
"role": "base_model",
"definition": "Model receiving the tokenizers.",
"value": {
"type": "concept_ref",
"versionId": "f89e740e-7204-4148-8934-85fbc73f323c",
"key": "qwen2_5_0_5b"
}
},
{
"role": "initialisation",
"definition": "Embedding initialisation of the new vocabularies.",
"value": {
"type": "concept_ref",
"versionId": "01087ca0-917c-46f8-a25e-ebc786018d72",
"key": "fvt_init"
}
},
{
"role": "procedure",
"definition": "Training of every arm.",
"value": {
"type": "concept_ref",
"versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
"key": "continued_pretraining"
}
},
{
"role": "budget",
"definition": "Continued-pretraining tokens per arm.",
"value": {
"type": "decimal",
"value": "1000000000",
"unit": "tokens"
}
},
{
"role": "reference_arm",
"definition": "Arm the residual is measured against.",
"value": {
"type": "text",
"value": "Qwen2.5-0.5B with its own tokenizer after the same continued pretraining"
}
},
{
"role": "metric",
"definition": "Quantity bounded.",
"value": {
"type": "concept",
"key": "residual_ratio"
}
},
{
"role": "texts",
"definition": "Texts pooled.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "formal_domains"
}
},
{
"role": "threshold",
"definition": "Upper bound on the metric.",
"value": {
"type": "decimal",
"value": "0.5",
"unit": "ratio"
}
},
{
"role": "cost_metric",
"definition": "Polish quantity whose increase is bounded.",
"value": {
"type": "concept_ref",
"versionId": "ed4b79d2-8435-4a8b-b3c2-f466d5084d7b",
"key": "pooled_polish_bpb"
}
},
{
"role": "cost_limit",
"definition": "Upper bound on the Polish increase.",
"value": {
"type": "decimal",
"value": "0.02",
"unit": "bits per byte"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →