Finding

Sign in with GitHub
← Publications

Finding · P265 · Author-curated

On 589 translation-paired MMLU items, continued pretraining of Qwen2.5-1.5B on 500M Polish-heavy tokens with its original tokenizer narrows its English-minus-Polish accuracy gap by 0.1036 while English accuracy changes by -0.1036 (95% interval -0.1409 to -0.0645) and Polish accuracy by 0.0000 (-0.0458 to 0.0424): erosion, not improved Polish access.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Classified by the interpretive rule

metric
English-minus-Polish accuracy gapQuantity whose change is classified.
benchmark
Translation-paired MMLU itemsItem pairs scored.
subject
Qwen2.5-1.5B continued with its own tokenizerModel after continued pretraining with its original tokenizer.
comparator
Qwen2.5-1.5BModel before continued pretraining.
gap change
-0.1036 accuracySubject minus comparator gap.
english change
-0.1036 accuracySubject minus comparator English accuracy.
polish change
0.0000 accuracySubject minus comparator Polish accuracy.
polish interval low
-0.0458 accuracyLower bound of the 95% interval of polish_change.
polish interval high
0.0424 accuracyUpper bound of the 95% interval of polish_change.
training
500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per armTraining the continued-pretraining arms received.
classification
erosion: the gap narrows while English accuracy falls and Polish accuracy is unchangedClass assigned by the rule.
rule
improved Polish access only if Polish accuracy rises with a 95% interval excluding zero; a narrowing carried by falling English accuracy with flat Polish accuracy is erosionInterpretive rule applied.

Experimental provenance

Method and evaluation protocol
Apply the plan's interpretive rule to the three paired bootstrap contrasts on MMLU.
Dataset
589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.Version: cais/mmlu@c30699e8356da336a370243923dbaf21066bb9fe; openGPT-X/mmlux@7fb62cf05c1ce9ea550dbfd2ea6c145204344882 · Access: public
Reported results
gap -0.1036, English -0.1036, Polish 0.0000.
Uncertainty and replication
Paired bootstrap 95% intervals; the English decline may include loss of memorised MMLU items (contamination decay), which the design does not separate from erosion.
Limitations
MMLU is the secondary benchmark; on Belebele, the primary, the gap does not narrow. One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Classified by the interpretive rule

The change of the metric from the comparator to the subject on the benchmark, with the English and Polish accuracy changes stated, receives the classification under the interpretive rule of the plan.

Key classified_by_rule · version 20eefc46-5718-45f8-bf30-405ada15f1c8

Translation-paired MMLU items

589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.

Key mmlu_paired_items · version 22e9d053-dd03-4cb2-ba2d-bea9e71a1260

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

English-minus-Polish accuracy gap

A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.

Key english_minus_polish_gap · version 5908110d-d39f-40bc-b50d-cdb85932e670

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

derived from

bb85a090-f854-4353-89d1-d74a1ee2ff1f

English accuracy change.

derived from

e97b0312-fa8b-4017-90e5-4b1acac3b019

Polish accuracy change.