Experiment · E13
Does scoring the translation-paired Belebele and MMLU items under every cyclic order of their options change any pre-registered decision about the English-minus-Polish accuracy gap of Qwen2.5-1.5B and its two continued-pretraining arms?
Qwen2.5-1.5B and its two arms after 500,170,752 Qwen tokens of continued pretraining (original tokenizer; APT4 by Fast Vocabulary Transfer) on the 598 Belebele and 589 MMLU translation pairs left after the frozen structural exclusions, in English and in Polish: the four cyclic orders of each item's options, of which the identity order is already scored, and a content-free input that measures each model's prior over option letters. The forward-discordant pair sets are re-derived from the new scores. Likelihood scoring only; no training.
Prerequisites and protocol
- Access needs
- The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer reference that the APT4 arm's loader checks against is gated on Hugging Face. Scoring runs on Apple silicon, 128 GB unified memory.
- Suggested protocol
- Commit a design, with its decision rule, before any scoring. Estimator: permutation-marginalised accuracy, in which each option's content is scored in every answer slot and the prediction is the option with the highest log-probability per continuation token averaged over the four orders; the standard deviation over orders is reported per cell. Pre-checks: per-slot prediction counts and accuracy under each permuted order; the share of items whose prediction is stable across orders, with a sensitivity analysis restricted to them. Decision rule: if a permutation-marginalised estimate moves any pre-registered bin of the gap, of its continued-pretraining or transplant contrasts, of the scaffold-swap bound or of the translation-assist classification, the affected findings are corrected with the permutation-marginalised endpoint primary and the identity-order endpoint as sensitivity. Admissible outcomes: the endpoint holds; a correction; or the arm cells carry too little signal above the 0.25 chance level to support the strength of the continued-pretraining contrast. Kill criterion: every permutation-marginalised estimate within 1 percentage point of its identity-order value. Letter or position priors are caught by dispersion across orders, which a prior-driven prediction cannot avoid; position effects introduced by the reordering itself are caught by the pre-checks. Scoring: scripts/02_probe_paired_mcq.py of experiments/E11-cross-language-access with an option-order argument added and the APT4 tokenizer loader unchanged; count the pairs that enter or leave the forward-discordant sets of experiments/E12-activation-patching.
Accepted plan
No accepted plan. Available work need not have a complete protocol or source commit.
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
No attempt registered. Work status and findings are independent of attempts.
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →