Experiment

Sign in with GitHub
← Experiments

Experiment · E14

How much of the English-minus-Polish accuracy gap of Qwen2.5-1.5B on translation-paired MMLU items, and of the contrasts with its continued-pretraining arms, rests on items with a wrong gold answer or culture-bound content?

Available · Proposed by @stw2 · Unassigned

The 589 MMLU translation pairs and the existing per-item scores of Qwen2.5-1.5B and its two continued-pretraining arms. Pairs flagged by published item-level annotations of MMLU ground-truth errors or of cultural sensitivity are dropped, and the gap and every arm contrast are re-estimated with the original estimators. No forward passes.

Prerequisites and protocol

Access needs
No restricted materials: the item files and per-item scores are public files of experiments/E11-cross-language-access; the annotation sources are pinned in the design.
Suggested protocol
Commit a design that names and pins the annotation sources before joining. Join the scored pairs to the annotations by subject and question, report the flagged share by flag type, drop the flagged pairs, and recompute gaps, contrasts and intervals with scripts/03_analyze.py of experiments/E11-cross-language-access; once the option-permutation audit exists, report its endpoint alongside. The existing exclusions test structural malformation only, not whether the gold answer is right. Kill criterion: fewer than 2% of pairs flagged and the gap unchanged within 1 percentage point.

Accepted plan

No accepted plan. Available work need not have a complete protocol or source commit.

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

No attempt registered. Work status and findings are independent of attempts.

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →