Experiment

Sign in with GitHub
← Experiments

Experiment · E20

Does the English-minus-Polish accuracy gap keep its sign and size when natively Polish items are machine-translated into English, so that translation falls on the English side?

Available · Proposed by @stw2 · Unassigned

The 300 LLMzSzŁ STEM and 250 PES questions of the reasoning-language experiment, written in Polish, machine-translated into English and scored in both languages by Qwen2.5-1.5B and its two continued-pretraining arms with the paired likelihood scorer under the option-permutation ensemble; per-item round-trip translation quality as a covariate and a seeded adequacy rating of 50 items. PES items have five options and LLMzSzŁ items four to six, so chance differs by item.

Prerequisites and protocol

Access needs
The LLMzSzŁ and PES item sets are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E03-reasoning-language at the pinned dataset revisions. The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. One machine-translation pass; scoring on Apple silicon, 128 GB unified memory.
Suggested protocol
Precondition: the option-permutation audit's endpoint. Commit a design, naming the translation system and its version, before translating. Score both languages with scripts/02_probe_paired_mcq.py of experiments/E11-cross-language-access extended to more than four options, and compare the gap with the same models' gaps on translation-paired Belebele and MMLU items under the same endpoint. Translation degradation follows the translated side: a gap produced by translation shrinks or reverses when the direction is reversed, and a model-side gap keeps its sign in both directions. Kill criterion: a mirror gap near zero or negative means most of the original gap is a direction artifact; a mirror gap near the original also yields a natively Polish paired item suite.

Accepted plan

No accepted plan. Available work need not have a complete protocol or source commit.

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

No attempt registered. Work status and findings are independent of attempts.

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →