Exact proposal revision
How much of the English-minus-Polish accuracy gap of Qwen2.5-1.5B on translation-paired MMLU items, and of the contrasts with its continued-pretraining arms, rests on items with a wrong gold answer or culture-bound content?
The 589 MMLU translation pairs and the existing per-item scores of Qwen2.5-1.5B and its two continued-pretraining arms. Pairs flagged by published item-level annotations of MMLU ground-truth errors or of cultural sensitivity are dropped, and the gap and every arm contrast are re-estimated with the original estimators. No forward passes.
Access and suggested protocol
- Access needs
- No restricted materials: the item files and per-item scores are public files of experiments/E11-cross-language-access; the annotation sources are pinned in the design.
- Suggested protocol
- Commit a design that names and pins the annotation sources before joining. Join the scored pairs to the annotations by subject and question, report the flagged share by flag type, drop the flagged pairs, and recompute gaps, contrasts and intervals with scripts/03_analyze.py of experiments/E11-cross-language-access; once the option-permutation audit exists, report its endpoint alongside. The existing exclusions test structural malformation only, not whether the gold answer is right. Kill criterion: fewer than 2% of pairs flagged and the gap unchanged within 1 percentage point.
Selected exact hypotheses and premises
Premise · P258
fc557219-826d-4b88-9b16-d041740450b8The continued-pretraining contrast re-estimated without flagged pairs.
Premise · P261
f86b68aa-76a6-4896-9b3c-6936af4638dbThe transplant contrast re-estimated without flagged pairs.
Premise · P283
b1699637-dcb7-40e3-836e-4de6c3d4eb01The translation-assist recovery on the same MMLU items.
Reason for this revision
Initial proposal.