Accepted plan

Sign in with GitHub
← Experiment E11

Immutable accepted plan · retrospective

Paired English and Polish likelihood multiple-choice scoring of Qwen2.5-1.5B and two continued-pretraining checkpoints on Belebele and MMLU

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 1: a data audit excludes 13 structurally malformed item pairs (2 Belebele, 11 MMLU) from the analysis, and the analysis script is added. Adopted after the full scoring run had started and the base model's cells had been scored, before the analysis ran.

Public source

Plan

Prediction
The base model's gap on Belebele is above zero; continued pretraining narrows the gap; the tokenizer transplant changes the gap by at most 3 percentage points; the MMLU gap differs between STEM and non-STEM subjects (exploratory).
Protocol
As the locked design, with Amendment 1: 03 drops the pairs listed in data/exclusions.json from every cell before computing gaps, sign tests, paired bootstrap contrasts (2000 resamples, seed 20260703), the decomposition, the STEM split and the floor rule; 04 writes flat headline values. The criterion is structural malformation (option collapse, gold answer absent or leaked, empty or duplicate options); translation style stays part of the Polish condition.
Dataset
598 Belebele and 589 MMLU item pairs after excluding 13 structurally malformed pairs from the 600 and 600 sampled; 111 MMLU pairs are STEM.
Split
Test splits only: a seeded sample of 600 items per benchmark (seed 42); every sampled pair is scored by every model in both languages.
Access needs
The two continued-pretraining checkpoints are restricted materials: ask the Room owner, or rebuild them with the control-arm experiment's scripts. The APT4 tokenizer that the arm B loader checks against is in a gated repository.
Configurations
Three models (Qwen2.5-1.5B; its checkpoint after 500M tokens of continued pretraining with the original tokenizer; its checkpoint with APT4 by FVT after the same continued pretraining) by two languages by two benchmarks. Scaffolds Answer (letter): for English and Odpowiedź (litera): for Polish. bf16 on Apple MPS. The transplant arm's tokenizer loads through the whitespace-canary loader of the control-arm experiment's second plan version.
Metric
English-minus-Polish likelihood multiple-choice accuracy gap per model (length-normalised primary, raw alongside), with a paired normal 95% interval and an exact sign test on discordant pairs; differences of gaps and of per-language accuracies between models with paired bootstrap 95% intervals.
Seeds
42 (item sampling); 20260703 (bootstrap, 2000 resamples); scoring uses no randomness
Interpretation rule
H1 is confirmed if the 95% interval of the base model's English-minus-Polish gap on Belebele lies above zero; gaps below 3, from 3 to 8 and above 8 percentage points are small, moderate and large. H2 is confirmed if the gap difference between the continued-pretraining checkpoint and the base model is below zero with a 95% interval excluding zero; improved Polish access is claimed only if Polish accuracy rises with an interval excluding zero, and a narrowing carried by falling English accuracy with flat Polish accuracy is erosion. H3: an absolute gap difference between the two checkpoints of at most 3 percentage points is the null result; a widening above 3 points with an interval excluding zero is a knowledge-side injury. H4 is exploratory with no bin. Any model, language and benchmark cell whose 95% interval for accuracy includes 0.25 demotes every contrast involving it to exploratory. Belebele is the primary benchmark.
Resources
Apple silicon, 128 GB unified memory; torch MPS bf16, scoring only; hours estimated.
Prior work
arXiv:2604.10799v1 reports Belebele and INCLUDE scores of the Bielik v3 models before and after the APT4 transplant; the openGPT-X Polish MMLU translation is described in arXiv:2410.08928.

Selected exact hypotheses and premises

Hypothesis · H30

50041e38-5826-4b54-b39b-fdff21892c61

Gap of the base model on Belebele.

Hypothesis · H31

5908110d-d39f-40bc-b50d-cdb85932e670

Effect of continued pretraining on the gap.

Hypothesis · H32

db9ee49a-78d0-476e-a9c0-9a737eac263a

Effect of the tokenizer transplant on the gap.

Hypothesis · H33

50390697-7a2b-44cd-bf3c-ffd92d092fa7

Exploratory STEM split on MMLU.

Premise · P40

15132b6d-26e7-49b9-8217-b144a28236c7

The preservation statement whose language-conditional component the gap measures.

Premise · P37

e0184bf2-dccb-469b-85be-cefff666bd50

Belebele scores of the Bielik 11B pair across European languages.

Premise · P35

cbcebd2f-77d5-48ae-9534-159d89911958

Multilingual examination scores of the Bielik 11B pair.

Premise · P113

faa76c5f-a834-41c9-8e3f-e4422de4c981

Continued pretraining with the original tokenizer lowers Polish bits per byte, so the arm scored here did learn Polish text.

Premise · P104

9a53082a-65e8-4c6a-82be-33dda5810584

At matched continued pretraining the transplant arm's Polish bits per byte is close to the original-tokenizer arm's.

Premise · P108

0c90c839-74ad-4f13-8f54-0b2dfbc9afbb

At matched continued pretraining the transplant arm's English bits per byte stays above the original-tokenizer arm's.

Premise · P134

4cd12925-1344-4a79-9c5b-4c2c52b50ce3

Polish examination multiple-choice accuracy of both arms stays near chance through 500M tokens.

Premise · P135

01b54d59-cc52-4829-8f63-7ec8298cbe0c

Polish medical examination multiple-choice accuracy of both arms is equal at 500M tokens and near chance.

Premise · P130

4ac9d8db-adb9-49c5-b0dd-e12ea68e549b

A likely memorised English benchmark text has the base model's lowest bits per byte.

Premise · P122

7ff16749-a07c-49e4-b96e-f230b7c9680c

Continued pretraining raises bits per byte on that memorised text, a decay that can lower English accuracy without any change in access.