Finding

Sign in with GitHub
← Publications

Finding · P261 · Author-curated

On 589 translation-paired MMLU items, after the same 500M tokens of Polish-heavy continued pretraining, the English-minus-Polish accuracy gap of Qwen2.5-1.5B is 0.0951 with APT4 by FVT and 0.1104 with its original tokenizer, a difference of -0.0153 (95% interval -0.0662 to 0.0357).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
English-minus-Polish accuracy gapQuantity compared.
benchmark
Translation-paired MMLU itemsItem pairs scored.
subject
Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continuedModel with APT4 by FVT after the same continued pretraining.
subject value
0.0951 accuracyMetric of the subject.
comparator
Qwen2.5-1.5B continued with its own tokenizerModel after continued pretraining with its original tokenizer.
comparator value
0.1104 accuracyMetric of the comparator.
value
-0.0153 accuracySubject minus comparator.
interval low
-0.0662 accuracyLower bound of the paired bootstrap 95% interval.
interval high
0.0357 accuracyUpper bound of the paired bootstrap 95% interval.
training
500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per armTraining the continued-pretraining arms received.

Experimental provenance

Method and evaluation protocol
Item-level paired bootstrap over the analysed pairs (2000 resamples, numpy PCG64 seed 20260703): every resample draws the same pairs for both models and both languages; the estimate is the difference on all pairs.
Dataset
589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.Version: cais/mmlu@c30699e8356da336a370243923dbaf21066bb9fe; openGPT-X/mmlux@7fb62cf05c1ce9ea550dbfd2ea6c145204344882 · Access: public
Reported results
0.0951 against 0.1104: -0.0153 [-0.0662, 0.0357].
Uncertainty and replication
Paired bootstrap percentile 95% interval; one training run per checkpoint, so training variance is not included.
Limitations
One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. The two models compared may favour different option letters, so the difference can reflect letter priors as well as item knowledge. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

Translation-paired MMLU items

589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.

Key mmlu_paired_items · version 22e9d053-dd03-4cb2-ba2d-bea9e71a1260

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

English-minus-Polish accuracy gap

A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.

Key english_minus_polish_gap · version 5908110d-d39f-40bc-b50d-cdb85932e670

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

derived from

3fffd423-ca29-41c5-b212-c6473f9cb468

Per-model gap and accuracies.

derived from

58bbfcef-0979-4d59-b6d1-533aac48d816

Per-model gap and accuracies.