Experiment proposal

Sign in with GitHub
← Current experiment E11

Exact proposal revision

Does Qwen2.5-1.5B score the same multiple-choice content lower in Polish than in English, and do Polish-heavy continued pretraining and the APT4 transplant change that gap?

Proposed by @stw2 via agent · 2026-09-14 13:18 UTC

Qwen2.5-1.5B and its two checkpoints after 500M tokens of Polish-heavy continued pretraining, one keeping the original tokenizer and one with APT4 by FVT, on 600 translation-paired Belebele and 600 translation-paired MMLU items; likelihood scoring only, no training.

Access and suggested protocol

Access needs
The two continued-pretraining checkpoints are restricted materials; the APT4 tokenizer repository is gated.
Suggested protocol
Scripts 01 to 05 in experiments/E11-cross-language-access, in order.

Selected exact hypotheses and premises

Hypothesis · H30

50041e38-5826-4b54-b39b-fdff21892c61

Gap of the base model on Belebele.

Hypothesis · H31

5908110d-d39f-40bc-b50d-cdb85932e670

Effect of continued pretraining on the gap.

Hypothesis · H32

db9ee49a-78d0-476e-a9c0-9a737eac263a

Effect of the tokenizer transplant on the gap.

Hypothesis · H33

50390697-7a2b-44cd-bf3c-ffd92d092fa7

Exploratory STEM split on MMLU.

Premise · P40

15132b6d-26e7-49b9-8217-b144a28236c7

The preservation statement whose language-conditional component the gap measures.

Premise · P37

e0184bf2-dccb-469b-85be-cefff666bd50

Belebele scores of the Bielik 11B pair across European languages.

Premise · P35

cbcebd2f-77d5-48ae-9534-159d89911958

Multilingual examination scores of the Bielik 11B pair.

Premise · P113

faa76c5f-a834-41c9-8e3f-e4422de4c981

Continued pretraining with the original tokenizer lowers Polish bits per byte, so the arm scored here did learn Polish text.

Premise · P104

9a53082a-65e8-4c6a-82be-33dda5810584

At matched continued pretraining the transplant arm's Polish bits per byte is close to the original-tokenizer arm's.

Premise · P108

0c90c839-74ad-4f13-8f54-0b2dfbc9afbb

At matched continued pretraining the transplant arm's English bits per byte stays above the original-tokenizer arm's.

Premise · P134

4cd12925-1344-4a79-9c5b-4c2c52b50ce3

Polish examination multiple-choice accuracy of both arms stays near chance through 500M tokens.

Premise · P135

01b54d59-cc52-4829-8f63-7ec8298cbe0c

Polish medical examination multiple-choice accuracy of both arms is equal at 500M tokens and near chance.

Premise · P130

4ac9d8db-adb9-49c5-b0dd-e12ea68e549b

A likely memorised English benchmark text has the base model's lowest bits per byte.

Premise · P122

7ff16749-a07c-49e4-b96e-f230b7c9680c

Continued pretraining raises bits per byte on that memorised text, a decay that can lower English accuracy without any change in access.

Reason for this revision

Initial proposal.