Experiment · E20
Does the English-minus-Polish accuracy gap keep its sign and size when natively Polish items are machine-translated into English, so that translation falls on the English side?
The 300 LLMzSzŁ STEM and 250 PES questions of the reasoning-language experiment, written in Polish, machine-translated into English and scored in both languages by Qwen2.5-1.5B and its two continued-pretraining arms with the paired likelihood scorer under the option-permutation ensemble; per-item round-trip translation quality as a covariate and a seeded adequacy rating of 50 items. PES items have five options and LLMzSzŁ items four to six, so chance differs by item.
Prerequisites and protocol
- Access needs
- The LLMzSzŁ and PES item sets are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E03-reasoning-language at the pinned dataset revisions. The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. One machine-translation pass; scoring on Apple silicon, 128 GB unified memory.
- Suggested protocol
- Precondition: the option-permutation audit's endpoint. Commit a design, naming the translation system and its version, before translating. Score both languages with scripts/02_probe_paired_mcq.py of experiments/E11-cross-language-access extended to more than four options, and compare the gap with the same models' gaps on translation-paired Belebele and MMLU items under the same endpoint. Translation degradation follows the translated side: a gap produced by translation shrinks or reverses when the direction is reversed, and a model-side gap keeps its sign in both directions. Kill criterion: a mirror gap near zero or negative means most of the original gap is a direction artifact; a mirror gap near the original also yields a natively Polish paired item suite.
Accepted plan
No accepted plan. Available work need not have a complete protocol or source commit.
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
No attempt registered. Work status and findings are independent of attempts.
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →