Exact proposal revision
Does Qwen2.5-1.5B score the same multiple-choice content lower in Polish than in English, and do Polish-heavy continued pretraining and the APT4 transplant change that gap?
Qwen2.5-1.5B and its two checkpoints after 500M tokens of Polish-heavy continued pretraining, one keeping the original tokenizer and one with APT4 by FVT, on 600 translation-paired Belebele and 600 translation-paired MMLU items; likelihood scoring only, no training.
Access and suggested protocol
- Access needs
- The two continued-pretraining checkpoints are restricted materials; the APT4 tokenizer repository is gated.
- Suggested protocol
- Scripts 01 to 05 in experiments/E11-cross-language-access, in order.
Selected exact hypotheses and premises
Premise · P40
15132b6d-26e7-49b9-8217-b144a28236c7The preservation statement whose language-conditional component the gap measures.
Premise · P37
e0184bf2-dccb-469b-85be-cefff666bd50Belebele scores of the Bielik 11B pair across European languages.
Premise · P35
cbcebd2f-77d5-48ae-9534-159d89911958Multilingual examination scores of the Bielik 11B pair.
Premise · P113
faa76c5f-a834-41c9-8e3f-e4422de4c981Continued pretraining with the original tokenizer lowers Polish bits per byte, so the arm scored here did learn Polish text.
Premise · P104
9a53082a-65e8-4c6a-82be-33dda5810584At matched continued pretraining the transplant arm's Polish bits per byte is close to the original-tokenizer arm's.
Premise · P108
0c90c839-74ad-4f13-8f54-0b2dfbc9afbbAt matched continued pretraining the transplant arm's English bits per byte stays above the original-tokenizer arm's.
Premise · P134
4cd12925-1344-4a79-9c5b-4c2c52b50ce3Polish examination multiple-choice accuracy of both arms stays near chance through 500M tokens.
Premise · P135
01b54d59-cc52-4829-8f63-7ec8298cbe0cPolish medical examination multiple-choice accuracy of both arms is equal at 500M tokens and near chance.
Premise · P130
4ac9d8db-adb9-49c5-b0dd-e12ea68e549bA likely memorised English benchmark text has the base model's lowest bits per byte.
Premise · P122
7ff16749-a07c-49e4-b96e-f230b7c9680cContinued pretraining raises bits per byte on that memorised text, a decay that can lower English accuracy without any change in access.
Reason for this revision
Initial proposal.