Finding

Sign in with GitHub
← Publications

Finding · P264 · Author-curated

After 500M tokens of Polish-heavy continued pretraining with its original tokenizer, Qwen2.5-1.5B's Polish likelihood multiple-choice accuracy rises on neither translation-paired Belebele items (difference -0.1171, 95% interval -0.1605 to -0.0702) nor translation-paired MMLU items (difference 0.0000, 95% interval -0.0458 to 0.0424).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Raised on a count of benchmarks

metric
Likelihood multiple-choice accuracyQuantity compared.
language
PolishLanguage of the items.
subject
Qwen2.5-1.5B continued with its own tokenizerModel after continued pretraining with its original tokenizer.
comparator
Qwen2.5-1.5BModel before continued pretraining.
count raised
0 benchmarksBenchmarks where the subject's Polish accuracy is higher with an interval above zero.
count total
2 benchmarksBenchmarks compared.
benchmarks
translation-paired Belebele items; translation-paired MMLU itemsBenchmarks compared.
training
500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per armTraining the continued-pretraining arms received.

Experimental provenance

Method and evaluation protocol
Count the benchmarks whose paired bootstrap 95% interval of the Polish accuracy difference lies above zero.
Dataset
Translation-paired Belebele and MMLU items after the exclusions.Version: unspecified · Access: public
Reported results
0 of 2 benchmarks.
Uncertainty and replication
Paired bootstrap 95% intervals per benchmark.
Limitations
One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. Continued pretraining was 500M tokens without annealing; longer training may differ. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Raised on a count of benchmarks

The subject's metric exceeds the comparator's with a 95% interval of the difference above zero on count_raised of count_total benchmarks.

Key raised_on_count · version 72f15189-34bd-468a-a7dd-44743ac5531b

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

derived from

b1fc6f1f-ef47-464e-b1e1-0f078e4af255

Polish accuracy difference on this benchmark.

derived from

e97b0312-fa8b-4017-90e5-4b1acac3b019

Polish accuracy difference on this benchmark.