Immutable accepted plan · prospective
Paired English and Polish likelihood multiple-choice scoring of Qwen2.5-1.5B and two continued-pretraining checkpoints on Belebele and MMLU
The locked design of 19 July 2026, committed before any model was scored.
Public source
https://github.com/stw2/tokenizer-science-tax @ 3c128813c117fd41b46d3a9949c2c8751134df27
Reference checked 2026-09-14 13:18 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- The base model's gap on Belebele is above zero; continued pretraining narrows the gap; the tokenizer transplant changes the gap by at most 3 percentage points; the MMLU gap differs between STEM and non-STEM subjects (exploratory).
- Protocol
- 01 rebuilds seeded 600-item samples of Belebele (English and Polish) and MMLU, pairs MMLU with the openGPT-X Polish translation, gates pairing and gold parity and writes the item files. 02 scores every item zero-shot by the log-probability of each option letter after a scaffold in the item's language: twice on 10 items per file as a determinism gate, then on all items for the base model and both checkpoints. The analysis computes per-model English-minus-Polish gaps with sign tests, paired bootstrap contrasts between models, the decomposition of the continued-pretraining contrast into English and Polish accuracy changes, the STEM split on MMLU and the floor rule.
- Dataset
- 600 Belebele test items in English and Polish (professionally translated, parallel) and 600 MMLU test items in English with their openGPT-X machine translation into Polish; 112 MMLU pairs are STEM.
- Split
- Test splits only: a seeded sample of 600 items per benchmark (seed 42); every sampled pair is scored by every model in both languages.
- Access needs
- The two continued-pretraining checkpoints are restricted materials: ask the Room owner, or rebuild them with the control-arm experiment's scripts. The APT4 tokenizer that the arm B loader checks against is in a gated repository.
- Configurations
- Three models (Qwen2.5-1.5B; its checkpoint after 500M tokens of continued pretraining with the original tokenizer; its checkpoint with APT4 by FVT after the same continued pretraining) by two languages by two benchmarks. Scaffolds Answer (letter): for English and Odpowiedź (litera): for Polish. bf16 on Apple MPS. The transplant arm's tokenizer loads through the whitespace-canary loader of the control-arm experiment's second plan version.
- Metric
- English-minus-Polish likelihood multiple-choice accuracy gap per model (length-normalised primary, raw alongside), with a paired normal 95% interval and an exact sign test on discordant pairs; differences of gaps and of per-language accuracies between models with paired bootstrap 95% intervals.
- Seeds
- 42 (item sampling); 20260703 (bootstrap, 2000 resamples); scoring uses no randomness
- Interpretation rule
- H1 is confirmed if the 95% interval of the base model's English-minus-Polish gap on Belebele lies above zero; gaps below 3, from 3 to 8 and above 8 percentage points are small, moderate and large. H2 is confirmed if the gap difference between the continued-pretraining checkpoint and the base model is below zero with a 95% interval excluding zero; improved Polish access is claimed only if Polish accuracy rises with an interval excluding zero, and a narrowing carried by falling English accuracy with flat Polish accuracy is erosion. H3: an absolute gap difference between the two checkpoints of at most 3 percentage points is the null result; a widening above 3 points with an interval excluding zero is a knowledge-side injury. H4 is exploratory with no bin. Any model, language and benchmark cell whose 95% interval for accuracy includes 0.25 demotes every contrast involving it to exploratory. Belebele is the primary benchmark.
- Resources
- Apple silicon, 128 GB unified memory; torch MPS bf16, scoring only; hours estimated.
- Prior work
- arXiv:2604.10799v1 reports Belebele and INCLUDE scores of the Bielik v3 models before and after the APT4 transplant; the openGPT-X Polish MMLU translation is described in arXiv:2410.08928.
Selected exact hypotheses and premises
Premise · P40
15132b6d-26e7-49b9-8217-b144a28236c7The preservation statement whose language-conditional component the gap measures.
Premise · P37
e0184bf2-dccb-469b-85be-cefff666bd50Belebele scores of the Bielik 11B pair across European languages.
Premise · P35
cbcebd2f-77d5-48ae-9534-159d89911958Multilingual examination scores of the Bielik 11B pair.
Premise · P113
faa76c5f-a834-41c9-8e3f-e4422de4c981Continued pretraining with the original tokenizer lowers Polish bits per byte, so the arm scored here did learn Polish text.
Premise · P104
9a53082a-65e8-4c6a-82be-33dda5810584At matched continued pretraining the transplant arm's Polish bits per byte is close to the original-tokenizer arm's.
Premise · P108
0c90c839-74ad-4f13-8f54-0b2dfbc9afbbAt matched continued pretraining the transplant arm's English bits per byte stays above the original-tokenizer arm's.
Premise · P134
4cd12925-1344-4a79-9c5b-4c2c52b50ce3Polish examination multiple-choice accuracy of both arms stays near chance through 500M tokens.
Premise · P135
01b54d59-cc52-4829-8f63-7ec8298cbe0cPolish medical examination multiple-choice accuracy of both arms is equal at 500M tokens and near chance.
Premise · P130
4ac9d8db-adb9-49c5-b0dd-e12ea68e549bA likely memorised English benchmark text has the base model's lowest bits per byte.
Premise · P122
7ff16749-a07c-49e4-b96e-f230b7c9680cContinued pretraining raises bits per byte on that memorised text, a decay that can lower English accuracy without any change in access.