Finding

Sign in with GitHub
← Publications

Finding · P84 · Author-curated

On 250 GSM8K-PL problems with Polish chain-of-thought, Bielik-PL-11B-v3.0-Instruct has accuracy 0.884 and Bielik-11B-v3.0-Instruct 0.856.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Answer accuracyQuantity compared.
benchmark
GSM8K-PL problemsItems the values are measured on.
language
PolishLanguage of the problems and the reasoning.
setting
Polish chain-of-thoughtEvaluation setting.
subject
Bielik-PL-11B-v3.0-InstructModel with the APT4 tokenizer.
subject value
0.884 proportionAccuracy of the subject.
comparator
Bielik-11B-v3.0-InstructModel with the Mistral-derived tokenizer.
comparator value
0.856 proportionAccuracy of the comparator.
scope
5-shot; greedy decoding; MLX bf16; max_tokens 1280Decoding.

Experimental provenance

Method and evaluation protocol
Fraction of problems whose final numeric answer equals the gold answer, per model, under Polish chain-of-thought.
Dataset
250 GSM8K-PL problems.Version: Eurolingua/gsm8kx 363d87617cd417b731dc860194a6935fd5ff49ae; openai/gsm8k 740312add88f781978c0658806c59bc2815b9866 · Access: public
Reported results
0.884 against 0.856.
Uncertainty and replication
Point values on 250 problems; no interval computed for this comparison.
Limitations
Machine-translated problems, few-shot prompts, greedy decoding with MLX bf16; not the paper's English GSM8K evaluation. Retrospective design: the GSM8K-PL token cap was raised to 1280 without a recorded amendment.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Values from results/analysis.json, accuracy[<model>|pl|gsm8k_pl].

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Answer accuracy

Fraction of questions whose extracted final answer equals the gold answer.

Key answer_accuracy · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Polish chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

Key polish_chain_of_thought · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

GSM8K-PL problems

250 GSM8K test problems in the Polish machine translation of gsm8kx, stratified by the digit length of the gold answer, with openai/gsm8k gold answers.

Key gsm8k_pl_problems · version 080dd7b3-7a21-419e-bf4e-c3770f5cdde6

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

related

36da68d9-97cf-4575-ad42-ed56e0e895e9

The same model pair on English GSM8K in the paper's evaluation; here machine-translated test problems with Polish reasoning.

related

1832c8a3-4857-4668-b703-73fdb0289e46

The statement the English GSM8K values bear on.