Exceeds by at least
At the layer where the model's metric on the evaluation_set most exceeds its metric on the comparison_set, both sets drawn from the benchmark under the scoring, the difference is at least margin.
Hypothesis
Sign in with GitHubHypothesis · H45 · Author-curated prediction
Qwen2.5-1.5B; MMLU translation pairs primary and Belebele translation pairs secondary; train and test split over items; every decoder layer; the sets are re-derived by the option-permutation audit. The layer of the maximum is not fixed in advance.
finding · premise
The scaffold-final residual patch on Qwen2.5-1.5B fails the pre-registered validity gate at all 3 layer pairs confirmed on 326 forward-discordant pairs, with breakage of 0.1600 at 25 into 25, 0.1600 at 25 into 23, 0.7400 at 1 into 27 against at most 0.05, so the patch does not classify the English-minus-Polish accuracy gap as routing-bound or representation-bound.Residual patching did not classify the gap as routing-bound or representation-bound.
finding · premise
On 589 translation-paired MMLU items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.5569 in English and 0.3430 in Polish, an English-minus-Polish gap of 0.2139 (95% interval 0.1672 to 0.2606).The English-minus-Polish gap on the MMLU translation pairs.
Loading research…
At the layer where the model's metric on the evaluation_set most exceeds its metric on the comparison_set, both sets drawn from the benchmark under the scoring, the difference is at least margin.
Translation pairs, after structural exclusions, that the model answers incorrectly in English and in Polish when each item's prediction is the option with the highest log-probability per continuation token averaged over the cyclic orders of the options.
Translation pairs, after structural exclusions, that the model answers correctly in English and incorrectly in Polish when each item's prediction is the option with the highest log-probability per continuation token averaged over the cyclic orders of the options.
Share of held-out items whose gold option gets the highest score from a linear read-out that was trained, on English passes of other items, to score from the residual-stream state at an option's final token whether that option is gold, and is applied at the same layer to the item's Polish pass.
Fraction of items for which the option whose content has the highest log-probability per continuation token, averaged over the cyclic orders of the options so that each content is scored in every answer slot, is the gold option.
589 pairs of MMLU test questions in English and in the openGPT-X machine translation into Polish: a seed-42 sample of 600 with the structurally malformed pairs excluded; 111 pairs belong to STEM subjects.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
{
"wording": "A linear read-out trained on English passes of Qwen2.5-1.5B and applied to its Polish passes selects the gold option of MMLU translation pairs that are forward-discordant under permutation-marginalised scoring at least 15 percentage points more often than on pairs wrong in both languages, at the best layer.",
"predicate": {
"type": "concept",
"key": "exceeds_by_at_least"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "cross_lingual_readout_accuracy"
}
},
{
"role": "model",
"definition": "Model read out.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "benchmark",
"definition": "Item pairs the sets are drawn from.",
"value": {
"type": "concept_ref",
"versionId": "22e9d053-dd03-4cb2-ba2d-bea9e71a1260",
"key": "mmlu_paired_items"
}
},
{
"role": "scoring",
"definition": "Scoring that defines the sets.",
"value": {
"type": "concept",
"key": "permutation_marginalised_mcq_accuracy"
}
},
{
"role": "evaluation_set",
"definition": "Pairs predicted to decode better.",
"value": {
"type": "concept",
"key": "permutation_discordant_pairs"
}
},
{
"role": "comparison_set",
"definition": "Pairs the evaluation set is compared with.",
"value": {
"type": "concept",
"key": "permutation_both_wrong_pairs"
}
},
{
"role": "margin",
"definition": "Lower bound on the difference.",
"value": {
"type": "decimal",
"value": "15",
"unit": "percentage points"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →