Finding

Sign in with GitHub
← Publications

Finding · P254 · Author-curated

On 598 translation-paired Belebele items, the English likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.6087 after 500M tokens of Polish-heavy continued pretraining with its original tokenizer and 0.7458 before it, a difference of -0.1371 (95% interval -0.1739 to -0.1020).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Likelihood multiple-choice accuracyQuantity compared.
language
EnglishLanguage of the items the accuracy is measured on.
benchmark
Translation-paired Belebele itemsItem pairs scored.
subject
Qwen2.5-1.5B continued with its own tokenizerModel after continued pretraining with its original tokenizer.
subject value
0.6087 accuracyMetric of the subject.
comparator
Qwen2.5-1.5BModel before continued pretraining.
comparator value
0.7458 accuracyMetric of the comparator.
value
-0.1371 accuracySubject minus comparator.
interval low
-0.1739 accuracyLower bound of the paired bootstrap 95% interval.
interval high
-0.1020 accuracyUpper bound of the paired bootstrap 95% interval.
training
500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per armTraining the continued-pretraining arms received.

Experimental provenance

Method and evaluation protocol
Item-level paired bootstrap over the analysed pairs (2000 resamples, numpy PCG64 seed 20260703): every resample draws the same pairs for both models and both languages; the estimate is the difference on all pairs.
Dataset
598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.Version: facebook/belebele@7899cdfa4e1e0d733fd77c848e2c273cb1d32be2 · Access: public
Reported results
0.6087 against 0.7458: -0.1371 [-0.1739, -0.1020].
Uncertainty and replication
Paired bootstrap percentile 95% interval; one training run per checkpoint, so training variance is not included.
Limitations
One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. The two models compared may favour different option letters, so the difference can reflect letter priors as well as item knowledge. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

Translation-paired Belebele items

598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.

Key belebele_paired_items · version e6ce775f-1927-4ab2-90fc-cdab7bcec624

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

derived from

74249c3b-d423-4eff-999e-6ad11900d3fc

Per-model gap and accuracies.

derived from

e6ce775f-1927-4ab2-90fc-cdab7bcec624

Per-model gap and accuracies.