Finding

Sign in with GitHub
← Publications

Finding · P72 · Author-curated

On 250 GSM8K test problems with five fixed worked examples and greedy decoding, Bielik-PL-11B-v3.0-Instruct scores 0.916 and Bielik-11B-v3.0-Instruct 0.928 (paired difference -0.012, 95% interval -0.036 to 0.012).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Exact answer accuracyQuantity compared.
benchmark
GSM8KBenchmark of the problems.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
subject value
0.916 share of problemsAccuracy of the subject.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
comparator value
0.928 share of problemsAccuracy of the comparator.
value
-0.012 difference in share of problemsPaired difference.
interval low
-0.036 difference in share of problemsLower bound of the 95% interval.
interval high
0.012 difference in share of problemsUpper bound of the 95% interval.
scope
250 test problems stratified by answer digit length (seed 20260703); five training problems as fixed examples; greedy decoding; at most 512 generated tokensProblems and protocol.

Experimental provenance

Method and evaluation protocol
Extract the answer after the last '####', or else the last number, strip commas and compare with the gold answer; paired bootstrap of the per-problem difference, 1,000 resamples, seed 20260703.
Dataset
GSM8K main configuration at revision 740312add88f781978c0658806c59bc2815b9866: 250 test problems and five training problems.Version: 740312add88f781978c0658806c59bc2815b9866 · Access: public
Reported results
0.916 against 0.928; difference -0.012 [-0.036, 0.012].
Uncertainty and replication
95% paired bootstrap interval over the problems, 1,000 resamples.
Limitations
Sign-only evidence by design: the interval covers zero. Fixed examples, greedy decoding and lenient answer extraction differ from the paper's evaluation, and both scores are higher than the paper's. The two models differ by 20B tokens of continued pretraining and by post-training as well as by the tokenizer; answer-only probes do not measure chain-of-thought arithmetic; MLX bf16 numerics differ from the paper's evaluation stack.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Table 6, p. 12, GSM8K column

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Exact answer accuracy

Share of problems whose extracted final answer equals the gold answer.

Key exact_answer_accuracy · version 59f84e17-2dbc-4e7a-b481-9fe8fc19f29b

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

GSM8K

Mathematical reasoning task of the English Open LLM Leaderboard.

Key gsm8k · version 36da68d9-97cf-4575-ad42-ed56e0e895e9

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

related

36da68d9-97cf-4575-ad42-ed56e0e895e9

Same sign as the reported difference, on a subset under a different protocol.