Gap recovery
Accuracy on the Polish twins with the intervention minus accuracy without it, divided by the English-minus-Polish accuracy gap.
Hypothesis
Sign in with GitHubHypothesis · H35 · Author-curated prediction
Qwen2.5-1.5B, likelihood-scored four-option prompts, per benchmark; no hooks.
finding · premise
On 598 translation-paired Belebele items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.7458 in English and 0.5485 in Polish, an English-minus-Polish gap of 0.1973 (95% interval 0.1558 to 0.2389).The English-minus-Polish accuracy gap of the base model on these translation pairs.
finding · premise
On 589 translation-paired MMLU items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.5569 in English and 0.3430 in Polish, an English-minus-Polish gap of 0.2139 (95% interval 0.1672 to 0.2606).The English-minus-Polish accuracy gap of the base model on these translation pairs.
Loading research…
Accuracy on the Polish twins with the intervention minus accuracy without it, divided by the English-minus-Polish accuracy gap.
Prepending the English twin's full prompt and one blank line to the Polish prompt, which keeps its Polish answer scaffold.
The share classifies each evaluation set as upper_class at or above upper_threshold, lower_class at or below lower_threshold and middle_class in between.
600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.
600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.
A model's likelihood multiple-choice accuracy on the English items minus its accuracy on the Polish versions of the same items, over the same item pairs.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
{
"wording": "The share of Qwen2.5-1.5B's English-minus-Polish accuracy gap on Belebele or MMLU translation pairs that is recovered by prepending the English twin's prompt classifies the gap as comprehension-dominant at 70% or more, deeper than comprehension at 30% or less, and mixed in between.",
"predicate": {
"type": "concept",
"key": "classified_by_share"
},
"roles": [
{
"role": "intervention",
"definition": "Change applied to each Polish prompt.",
"value": {
"type": "concept",
"key": "translation_assist"
}
},
{
"role": "share",
"definition": "Quantity that decides the class.",
"value": {
"type": "concept",
"key": "gap_recovery"
}
},
{
"role": "gap",
"definition": "Quantity the share is taken of.",
"value": {
"type": "concept_ref",
"versionId": "5908110d-d39f-40bc-b50d-cdb85932e670",
"key": "english_minus_polish_gap"
}
},
{
"role": "model",
"definition": "Model scored.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "evaluation_set_1",
"definition": "First set of items, classified on its own.",
"value": {
"type": "concept_ref",
"versionId": "cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea",
"key": "belebele_translation_pairs"
}
},
{
"role": "evaluation_set_2",
"definition": "Second set of items, classified on its own.",
"value": {
"type": "concept_ref",
"versionId": "cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea",
"key": "mmlu_translation_pairs"
}
},
{
"role": "upper_threshold",
"definition": "Share at or above which the upper class applies.",
"value": {
"type": "decimal",
"value": "70",
"unit": "percent"
}
},
{
"role": "upper_class",
"definition": "Class at or above upper_threshold.",
"value": {
"type": "text",
"value": "comprehension-dominant"
}
},
{
"role": "lower_threshold",
"definition": "Share at or below which the lower class applies.",
"value": {
"type": "decimal",
"value": "30",
"unit": "percent"
}
},
{
"role": "lower_class",
"definition": "Class at or below lower_threshold.",
"value": {
"type": "text",
"value": "deeper than comprehension"
}
},
{
"role": "middle_class",
"definition": "Class between the thresholds.",
"value": {
"type": "text",
"value": "mixed"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →