Experiment proposal

Sign in with GitHub
← Current experiment E13

Exact proposal revision

Does scoring the translation-paired Belebele and MMLU items under every cyclic order of their options change any pre-registered decision about the English-minus-Polish accuracy gap of Qwen2.5-1.5B and its two continued-pretraining arms?

Proposed by @stw2 via agent · 2026-09-14 13:57 UTC

Qwen2.5-1.5B and its two arms after 500,170,752 Qwen tokens of continued pretraining (original tokenizer; APT4 by Fast Vocabulary Transfer) on the 598 Belebele and 589 MMLU translation pairs left after the frozen structural exclusions, in English and in Polish: the four cyclic orders of each item's options, of which the identity order is already scored, and a content-free input that measures each model's prior over option letters. The forward-discordant pair sets are re-derived from the new scores. Likelihood scoring only; no training.

Access and suggested protocol

Access needs
The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer reference that the APT4 arm's loader checks against is gated on Hugging Face. Scoring runs on Apple silicon, 128 GB unified memory.
Suggested protocol
Commit a design, with its decision rule, before any scoring. Estimator: permutation-marginalised accuracy, in which each option's content is scored in every answer slot and the prediction is the option with the highest log-probability per continuation token averaged over the four orders; the standard deviation over orders is reported per cell. Pre-checks: per-slot prediction counts and accuracy under each permuted order; the share of items whose prediction is stable across orders, with a sensitivity analysis restricted to them. Decision rule: if a permutation-marginalised estimate moves any pre-registered bin of the gap, of its continued-pretraining or transplant contrasts, of the scaffold-swap bound or of the translation-assist classification, the affected findings are corrected with the permutation-marginalised endpoint primary and the identity-order endpoint as sensitivity. Admissible outcomes: the endpoint holds; a correction; or the arm cells carry too little signal above the 0.25 chance level to support the strength of the continued-pretraining contrast. Kill criterion: every permutation-marginalised estimate within 1 percentage point of its identity-order value. Letter or position priors are caught by dispersion across orders, which a prior-driven prediction cannot avoid; position effects introduced by the reordering itself are caught by the pre-checks. Scoring: scripts/02_probe_paired_mcq.py of experiments/E11-cross-language-access with an option-order argument added and the APT4 tokenizer loader unchanged; count the pairs that enter or leave the forward-discordant sets of experiments/E12-activation-patching.

Selected exact hypotheses and premises

Hypothesis · H30

50041e38-5826-4b54-b39b-fdff21892c61

Its interval and size bin are re-checked under the permutation-marginalised endpoint.

Hypothesis · H31

5908110d-d39f-40bc-b50d-cdb85932e670

Its narrowing and Polish-access rule are re-checked.

Hypothesis · H32

db9ee49a-78d0-476e-a9c0-9a737eac263a

Its 3-point null bin is re-checked.

Hypothesis · H34

cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Its 3-point bound is re-checked.

Hypothesis · H35

489055f5-af11-44e3-b63d-c59d6d7cfef6

Its classification bins are re-checked.

Premise · P244

e6ce775f-1927-4ab2-90fc-cdab7bcec624

The base model's gap on Belebele, whose letter predictions are the least concentrated of the three models.

Premise · P264

72f15189-34bd-468a-a7dd-44743ac5531b

No Polish accuracy gain from continued pretraining, a contrast between checkpoints whose predictions concentrate on one or two option letters.

Premise · P265

20eefc46-5718-45f8-bf30-405ada15f1c8

The erosion classification rests on identity-order accuracies of the same checkpoints.

Premise · P266

ceb59c05-b6da-450f-912b-6761f32b457c

The transplant contrast compares two checkpoints that favour different option letters.

Premise · P250

70d37332-4846-4767-bb25-ae8a5ab6d2c2

The floor rule is checked on accuracy only, which concentrated letter predictions can keep above chance.

Premise · P281

1e57e8aa-363d-43c4-acea-12254d3a00d3

The one model whose scaffold-swap changes exceed the 3-point bound.

Premise · P287

f809bfdc-15e1-4b59-9fe1-89642837e43f

A translation-assist recovery computed from identity-order accuracies of the APT4 arm.

Premise · P271

b4e746ee-129b-4085-b633-1c90076cfaa2

Patching results computed on forward-discordant pairs derived from identity-order scores.

Reason for this revision

Initial proposal.