Finding

Sign in with GitHub
← Publications

Finding · P257 · Author-curated

On 598 translation-paired Belebele items, after the same 500M tokens of Polish-heavy continued pretraining, the English likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.6104 with APT4 by FVT and 0.6087 with its original tokenizer, a difference of 0.0017 (95% interval -0.0385 to 0.0418).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Likelihood multiple-choice accuracyQuantity compared.
language
EnglishLanguage of the items the accuracy is measured on.
benchmark
Translation-paired Belebele itemsItem pairs scored.
subject
Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continuedModel with APT4 by FVT after the same continued pretraining.
subject value
0.6104 accuracyMetric of the subject.
comparator
Qwen2.5-1.5B continued with its own tokenizerModel after continued pretraining with its original tokenizer.
comparator value
0.6087 accuracyMetric of the comparator.
value
0.0017 accuracySubject minus comparator.
interval low
-0.0385 accuracyLower bound of the paired bootstrap 95% interval.
interval high
0.0418 accuracyUpper bound of the paired bootstrap 95% interval.
training
500,170,752 Qwen tokens of continued pretraining of each arm on the same document sequence, constant learning rate before any decay anneal, one run per armTraining the continued-pretraining arms received.

Experimental provenance

Method and evaluation protocol
Item-level paired bootstrap over the analysed pairs (2000 resamples, numpy PCG64 seed 20260703): every resample draws the same pairs for both models and both languages; the estimate is the difference on all pairs.
Dataset
598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.Version: facebook/belebele@7899cdfa4e1e0d733fd77c848e2c273cb1d32be2 · Access: public
Reported results
0.6104 against 0.6087: 0.0017 [-0.0385, 0.0418].
Uncertainty and replication
Paired bootstrap percentile 95% interval; one training run per checkpoint, so training variance is not included.
Limitations
One model family at 1.5B parameters; the checkpoints are pre-anneal, from one training run each. English MMLU items may be in Qwen2.5 pretraining data, so MMLU gaps can be inflated; Polish MMLU items are machine translations. The 13 excluded pairs were chosen after the base model's cells had been scored. The two models compared may favour different option letters, so the difference can reflect letter priors as well as item knowledge. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

Translation-paired Belebele items

598 pairs of the same FLORES passage, question and four options in English and in Polish from the Belebele test split: a seed-42 sample of 600 with the structurally malformed pairs excluded.

Key belebele_paired_items · version e6ce775f-1927-4ab2-90fc-cdab7bcec624

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

derived from

529861fe-27f7-4e88-b1fb-15b1f7327858

Per-model gap and accuracies.

derived from

74249c3b-d423-4eff-999e-6ad11900d3fc

Per-model gap and accuracies.