Narrows
The subject's value of the metric is smaller than the comparator's.
Hypothesis
Sign in with GitHubHypothesis · H31 · Author-curated prediction
Belebele is the primary benchmark and MMLU, with machine-translated Polish items, the secondary. A narrowing counts as improved Polish access only if Polish accuracy rises with a 95% interval excluding zero; a narrowing carried by falling English accuracy with flat Polish accuracy is erosion.
cited claim · premise
Across nine Polish and multilingual benchmarks, the Bielik v3 PL models closely preserve the performance of their original-tokenizer counterparts.Performance on Polish and multilingual benchmarks is stated to be preserved after vocabulary adaptation with continued pretraining.
cited claim · premise
On INCLUDE-base-44, averaged over 20 European languages, Bielik-PL-11B-v3.0-Instruct scores 53.92 and Bielik-11B-v3.0-Instruct 64.8.A multilingual benchmark average falls after the transplant with continued pretraining.
finding · premise
Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens lowers bits per byte on the Polish FineWeb2-HQ holdout from 1.2043 to 0.9345.Continued pretraining with the original tokenizer lowers Polish bits per byte, so the arm scored here did learn Polish text.
finding · premise
Untouched Qwen2.5-1.5B has 0.0524 bits per byte on GSM8K problems, the lowest of 10 texts scored; the next lowest is 0.3417 on Python code.A likely memorised English benchmark text has the base model's lowest bits per byte.
finding · premise
Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens raises bits per byte on GSM8K problems from 0.0524 to 0.4966.Continued pretraining raises bits per byte on that memorised text, a decay that can lower English accuracy without any change in access.
Loading research…
The subject's value of the metric is smaller than the comparator's.
A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
Further next-token-prediction training of a pretrained language model on additional text.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
{
"wording": "Continued pretraining of Qwen2.5-1.5B on 500M tokens of an 80/20 Polish-English mix narrows its English-minus-Polish likelihood multiple-choice accuracy gap on translation-paired items.",
"predicate": {
"type": "concept",
"key": "narrows"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "english_minus_polish_gap"
}
},
{
"role": "subject",
"definition": "Model after the intervention.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "comparator",
"definition": "Model before the intervention.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "intervention",
"definition": "Training applied to the comparator to obtain the subject.",
"value": {
"type": "concept_ref",
"versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
"key": "continued_pretraining"
}
},
{
"role": "training",
"definition": "Training the continued-pretraining arms received.",
"value": {
"type": "text",
"value": "500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per arm"
}
},
{
"role": "benchmarks",
"definition": "Item sets the gap is measured on.",
"value": {
"type": "text",
"value": "translation-paired Belebele items (primary); translation-paired MMLU items (secondary)"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →