Immutable accepted plan · retrospective
Paired English and Polish likelihood multiple-choice scoring of Qwen2.5-1.5B and two continued-pretraining checkpoints on Belebele and MMLU
Amendment 1: a data audit excludes 13 structurally malformed item pairs (2 Belebele, 11 MMLU) from the analysis, and the analysis script is added. Adopted after the full scoring run had started and the base model's cells had been scored, before the analysis ran.
Public source
https://github.com/stw2/tokenizer-science-tax @ a544bf7f4ee73f31ddaa43cf5b1218457461d94e
Reference checked 2026-09-14 13:18 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- The base model's gap on Belebele is above zero; continued pretraining narrows the gap; the tokenizer transplant changes the gap by at most 3 percentage points; the MMLU gap differs between STEM and non-STEM subjects (exploratory).
- Protocol
- As the locked design, with Amendment 1: 03 drops the pairs listed in data/exclusions.json from every cell before computing gaps, sign tests, paired bootstrap contrasts (2000 resamples, seed 20260703), the decomposition, the STEM split and the floor rule; 04 writes flat headline values. The criterion is structural malformation (option collapse, gold answer absent or leaked, empty or duplicate options); translation style stays part of the Polish condition.
- Dataset
- 598 Belebele and 589 MMLU item pairs after excluding 13 structurally malformed pairs from the 600 and 600 sampled; 111 MMLU pairs are STEM.
- Split
- Test splits only: a seeded sample of 600 items per benchmark (seed 42); every sampled pair is scored by every model in both languages.
- Access needs
- The two continued-pretraining checkpoints are restricted materials: ask the Room owner, or rebuild them with the control-arm experiment's scripts. The APT4 tokenizer that the arm B loader checks against is in a gated repository.
- Configurations
- Three models (Qwen2.5-1.5B; its checkpoint after 500M tokens of continued pretraining with the original tokenizer; its checkpoint with APT4 by FVT after the same continued pretraining) by two languages by two benchmarks. Scaffolds Answer (letter): for English and Odpowiedź (litera): for Polish. bf16 on Apple MPS. The transplant arm's tokenizer loads through the whitespace-canary loader of the control-arm experiment's second plan version.
- Metric
- English-minus-Polish likelihood multiple-choice accuracy gap per model (length-normalised primary, raw alongside), with a paired normal 95% interval and an exact sign test on discordant pairs; differences of gaps and of per-language accuracies between models with paired bootstrap 95% intervals.
- Seeds
- 42 (item sampling); 20260703 (bootstrap, 2000 resamples); scoring uses no randomness
- Interpretation rule
- H1 is confirmed if the 95% interval of the base model's English-minus-Polish gap on Belebele lies above zero; gaps below 3, from 3 to 8 and above 8 percentage points are small, moderate and large. H2 is confirmed if the gap difference between the continued-pretraining checkpoint and the base model is below zero with a 95% interval excluding zero; improved Polish access is claimed only if Polish accuracy rises with an interval excluding zero, and a narrowing carried by falling English accuracy with flat Polish accuracy is erosion. H3: an absolute gap difference between the two checkpoints of at most 3 percentage points is the null result; a widening above 3 points with an interval excluding zero is a knowledge-side injury. H4 is exploratory with no bin. Any model, language and benchmark cell whose 95% interval for accuracy includes 0.25 demotes every contrast involving it to exploratory. Belebele is the primary benchmark.
- Resources
- Apple silicon, 128 GB unified memory; torch MPS bf16, scoring only; hours estimated.
- Prior work
- arXiv:2604.10799v1 reports Belebele and INCLUDE scores of the Bielik v3 models before and after the APT4 transplant; the openGPT-X Polish MMLU translation is described in arXiv:2410.08928.
Selected exact hypotheses and premises
Premise · P40
15132b6d-26e7-49b9-8217-b144a28236c7The preservation statement whose language-conditional component the gap measures.
Premise · P37
e0184bf2-dccb-469b-85be-cefff666bd50Belebele scores of the Bielik 11B pair across European languages.
Premise · P35
cbcebd2f-77d5-48ae-9534-159d89911958Multilingual examination scores of the Bielik 11B pair.
Premise · P113
faa76c5f-a834-41c9-8e3f-e4422de4c981Continued pretraining with the original tokenizer lowers Polish bits per byte, so the arm scored here did learn Polish text.
Premise · P104
9a53082a-65e8-4c6a-82be-33dda5810584At matched continued pretraining the transplant arm's Polish bits per byte is close to the original-tokenizer arm's.
Premise · P108
0c90c839-74ad-4f13-8f54-0b2dfbc9afbbAt matched continued pretraining the transplant arm's English bits per byte stays above the original-tokenizer arm's.
Premise · P134
4cd12925-1344-4a79-9c5b-4c2c52b50ce3Polish examination multiple-choice accuracy of both arms stays near chance through 500M tokens.
Premise · P135
01b54d59-cc52-4829-8f63-7ec8298cbe0cPolish medical examination multiple-choice accuracy of both arms is equal at 500M tokens and near chance.
Premise · P130
4ac9d8db-adb9-49c5-b0dd-e12ea68e549bA likely memorised English benchmark text has the base model's lowest bits per byte.
Premise · P122
7ff16749-a07c-49e4-b96e-f230b7c9680cContinued pretraining raises bits per byte on that memorised text, a decay that can lower English accuracy without any change in access.