Higher in a language
The subject's metric on the benchmark items in language is higher than on the same items in comparison_language.
Hypothesis
Sign in with GitHubHypothesis · H30 · Author-curated prediction
Belebele test items with the same passage, question and options in English and Polish, each pair scored by the same model; comparison per item pair.
cited claim · premise
On Belebele Polish, Bielik-PL-11B-v3.0-Instruct scores 81.22 and Bielik-11B-v3.0-Instruct 82.11.Belebele measures Polish reading comprehension of the Bielik v3 models.
cited claim · premise
On Belebele, averaged over 28 European language variants, Bielik-PL-11B-v3.0-Instruct scores 77.41 and Bielik-11B-v3.0-Instruct 82.98.Belebele scores across European languages before and after the transplant.
Loading research…
The subject's metric on the benchmark items in language is higher than on the same items in comparison_language.
Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.
Multilingual multiple-choice reading comprehension benchmark on FLORES-200 passages.
The Polish language.
The English language.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
{
"wording": "Qwen2.5-1.5B's likelihood multiple-choice accuracy on translation-paired Belebele items is higher in English than in Polish.",
"predicate": {
"type": "concept",
"key": "higher_in_language"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "paired_mcq_accuracy"
}
},
{
"role": "subject",
"definition": "Model scored.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "benchmark",
"definition": "Items scored in both languages.",
"value": {
"type": "concept_ref",
"versionId": "e0184bf2-dccb-469b-85be-cefff666bd50",
"key": "belebele"
}
},
{
"role": "language",
"definition": "Language predicted to score higher.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "english"
}
},
{
"role": "comparison_language",
"definition": "Language of the same items predicted to score lower.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "polish"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →