English chain-of-thought gain
Answer accuracy of one model when its system prompt, few-shot reasoning traces and answer scaffold are in English, minus its answer accuracy when they are in Polish, on the same questions posed in Polish.
Hypothesis
Sign in with GitHubHypothesis · H8 · Author-curated prediction
Two 11B instruction-tuned Bielik v3 models that differ in tokenizer; questions posed in Polish; the language of the reasoning is set by the prompt.
finding · premise
APT4's tokenization contributes to arithmetic accuracy differences between Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct through its handling of the no-break space, not through digit segmentation.Digit segmentation, as measured, leaves the arithmetic gap only partly explained.
cited claim · premise
On GSM8K, Bielik-PL-11B-v3.0-Instruct scores 80.97 and Bielik-11B-v3.0-Instruct 85.60.The APT4 model scores lower on GSM8K than its original-tokenizer counterpart.
cited claim · premise
English-language capabilities of the Bielik v3 PL models remain largely intact.The statement the GSM8K values bear on.
cited claim · premise
On the Polish Medical Leaderboard, Bielik-PL-11B-v3.0-Instruct scores 48.42 percent and Bielik-11B-v3.0-Instruct 50.21 percent.The APT4 model scores lower on the PES-based benchmark than its counterpart.
Loading research…
Answer accuracy of one model when its system prompt, few-shot reasoning traces and answer scaffold are in English, minus its answer accuracy when they are in Polish, on the same questions posed in Polish.
The metric is larger for the subject than for the comparator on the evaluation items, in the language and setting stated.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The Polish language.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
{
"wording": "On Polish STEM questions, the accuracy gain from English over Polish chain-of-thought is larger for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct.",
"predicate": {
"type": "concept",
"key": "larger_for_subject"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "english_cot_gain"
}
},
{
"role": "subject",
"definition": "Model predicted to have the larger value.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "comparator",
"definition": "Model compared with.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "evaluation_items",
"definition": "Questions the metric is measured on.",
"value": {
"type": "text",
"value": "Polish STEM questions: machine-translated GSM8K problems, LLMzSzŁ mathematics, physics, science and biology examination questions, and PES medical specialization examination questions"
}
},
{
"role": "language",
"definition": "Language of the questions.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "polish"
}
},
{
"role": "setting",
"definition": "Evaluation setting.",
"value": {
"type": "text",
"value": "few-shot prompts; greedy decoding"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →