Differs by at most
The absolute difference between the subject's and the comparator's value of the metric is at most bound.
Hypothesis
Sign in with GitHubHypothesis · H32 · Author-curated prediction
Both checkpoints after the same 954 optimizer steps on the same document sequence. A widening above 3 percentage points whose 95% interval excludes zero would be a knowledge-side injury of the transplant.
cited claim · premise
Across nine Polish and multilingual benchmarks, the Bielik v3 PL models closely preserve the performance of their original-tokenizer counterparts.The transplant is stated to preserve Polish and multilingual benchmark performance.
cited claim · premise
On Belebele Polish, Bielik-PL-11B-v3.0-Instruct scores 81.22 and Bielik-11B-v3.0-Instruct 82.11.Polish Belebele scores of the 11B pair with and without the transplant.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on the Polish FineWeb2-HQ holdout is 0.9492 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.9345 for Qwen2.5-1.5B with its own tokenizer.At matched continued pretraining the transplant arm's Polish bits per byte is close to the original-tokenizer arm's.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on the English SlimPajama holdout is 0.9868 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.8540 for Qwen2.5-1.5B with its own tokenizer.At matched continued pretraining the transplant arm's English bits per byte stays above the original-tokenizer arm's.
finding · premise
On LLMzSzŁ STEM questions, likelihood accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer is 0.2333, 0.2933, 0.2967, 0.3100, 0.2833 and 0.2933, and of Qwen2.5-1.5B with its own tokenizer 0.2900, 0.3100, 0.3000, 0.2833, 0.2767 and 0.2800, after 0, 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.Polish examination multiple-choice accuracy of both arms stays near chance through 500M tokens.
finding · premise
On PES examination questions, likelihood accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer is 0.2280, 0.2320, 0.2320, 0.2200, 0.2280 and 0.2560, and of Qwen2.5-1.5B with its own tokenizer 0.2480, 0.2720, 0.2880, 0.2720, 0.2640 and 0.2560, after 0, 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.Polish medical examination multiple-choice accuracy of both arms is equal at 500M tokens and near chance.
Loading research…
The absolute difference between the subject's and the comparator's value of the metric is at most bound.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
Exchange of a pretrained model's tokenizer and vocabulary for those of another tokenizer, including initialisation of the new embeddings, before any further training.
A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
Polish-optimised tokenizer of the Bielik v3 PL models.
{
"wording": "At matched continued pretraining, the English-minus-Polish likelihood multiple-choice accuracy gap of Qwen2.5-1.5B with APT4 initialised by FVT differs by at most 3 percentage points from that of Qwen2.5-1.5B with its original tokenizer.",
"predicate": {
"type": "concept",
"key": "differs_by_at_most"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept_ref",
"versionId": "5908110d-d39f-40bc-b50d-cdb85932e670",
"key": "english_minus_polish_gap"
}
},
{
"role": "subject",
"definition": "Model with the replaced tokenizer.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "comparator",
"definition": "Model that kept the original tokenizer after the same continued pretraining.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "bound",
"definition": "Largest absolute difference predicted.",
"value": {
"type": "decimal",
"value": "0.03",
"unit": "accuracy"
}
},
{
"role": "tokenizer",
"definition": "Tokenizer of the subject.",
"value": {
"type": "concept_ref",
"versionId": "d497f94d-5373-4652-887c-55c001b6472c",
"key": "apt4"
}
},
{
"role": "intervention",
"definition": "Change that distinguishes the subject from the comparator.",
"value": {
"type": "concept_ref",
"versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
"key": "tokenizer_replacement"
}
},
{
"role": "training",
"definition": "Training the continued-pretraining arms received.",
"value": {
"type": "text",
"value": "500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per arm"
}
},
{
"role": "benchmarks",
"definition": "Item sets the gap is measured on.",
"value": {
"type": "text",
"value": "translation-paired Belebele items (primary); translation-paired MMLU items (secondary)"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →