Finding

Sign in with GitHub
← Publications

Finding · P82 · Author-curated

On 800 Polish STEM questions, the English chain-of-thought gain of Bielik-PL-11B-v3.0-Instruct is -0.0138 (accuracy 0.7488 with English and 0.7625 with Polish chain-of-thought), with a 95% bootstrap interval from -0.04 to 0.0125 and McNemar p 0.3382.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Estimate with interval

metric
English chain-of-thought gainQuantity estimated.
subject
Bielik-PL-11B-v3.0-InstructModel measured.
evaluation items
Polish STEM question setQuestions measured.
items
800 questionsNumber of questions.
value
-0.0138 proportionPoint estimate.
interval low
-0.04 proportionLower bound of the 95% interval.
interval high
0.0125 proportionUpper bound of the 95% interval.
p value
0.3382 p-valueTwo-sided p-value.
accuracy english cot
0.7488 proportionAccuracy under English chain-of-thought.
accuracy polish cot
0.7625 proportionAccuracy under Polish chain-of-thought.
setting
stratified paired item bootstrap, 2000 resamples, seed 20260703; exact McNemar test on 60 questions correct only with Polish and 49 only with English chain-of-thoughtHow the estimate was computed.

Experimental provenance

Method and evaluation protocol
English minus Polish chain-of-thought accuracy of one model over the same questions; interval from the bootstrap, p-value from the exact McNemar test on discordant questions.
Dataset
800 Polish STEM questions: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM and 250 PES examination questions.Version: Eurolingua/gsm8kx 363d87617cd417b731dc860194a6935fd5ff49ae; openai/gsm8k 740312add88f781978c0658806c59bc2815b9866; amu-cai/llmzszl-dataset 9c11d52a415147f1e2c44038a9d4d4155f064e1a; speakleash/PES-2018-2022 79ae158ceab8e4e083c8cfcacde56669407a94a2 · Access: restricted
Reported results
Gain -0.0138 [-0.04, 0.0125], McNemar p 0.3382; accuracy 0.7488 English, 0.7625 Polish.
Uncertainty and replication
95% stratified paired item bootstrap, 2000 resamples, seed 20260703; exact two-sided McNemar test.
Limitations
Retrospective design: the GSM8K-PL token cap was raised and capped answers were regenerated at 1280 tokens without a recorded amendment, the latter after an interim grading. MLX bf16 greedy decoding, compared within one numerical stack; GSM8K-PL is machine-translated.

Author’s note

Values from results/analysis.json, gain_en_over_pl and accuracy.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Estimate with interval

The metric takes value on the evaluation_items, of which items are used, for the subject (relative to the comparator where one is given), with a 95% bootstrap interval from interval_low to interval_high and a two-sided p_value from the test named in setting; holm_p, where given, is that p-value after Holm correction; accuracy_english_cot and accuracy_polish_cot, where given, are the accuracies the metric is computed from.

Key estimate_with_interval · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Polish STEM question set

800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.

Key polish_stem_questions · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

English chain-of-thought gain

Answer accuracy of one model when its system prompt, few-shot reasoning traces and answer scaffold are in English, minus its answer accuracy when they are in Polish, on the same questions posed in Polish.

Key english_cot_gain · version 0a102e08-95f9-4674-840b-802a26c1f2d6

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04