Finding

Sign in with GitHub
← Publications

Finding · P79 · Author-curated

On 250 PES examination questions, the English chain-of-thought gain of Bielik-PL-11B-v3.0-Instruct minus that of Bielik-11B-v3.0-Instruct is -0.036, with a 95% bootstrap interval from -0.116 to 0.044 (Holm-adjusted p 0.778).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Estimate with interval

subject
Bielik-PL-11B-v3.0-InstructModel whose gain is taken first.
comparator
Bielik-11B-v3.0-InstructModel whose gain is subtracted.
evaluation items
PES examination questionsQuestions measured.
items
250 questionsNumber of questions.
value
-0.036 proportionPoint estimate.
interval low
-0.116 proportionLower bound of the 95% interval.
interval high
0.044 proportionUpper bound of the 95% interval.
p value
0.389 p-valueTwo-sided p-value.
holm p
0.778 p-valuep-value after Holm correction over three benchmarks.
setting
paired item bootstrap, 2000 resamples, seed 20260703; greedy decoding, MLX bf16How the estimate was computed.

Experimental provenance

Method and evaluation protocol
As the pooled difference, restricted to one benchmark.
Dataset
250 PES examination questions.Version: speakleash/PES-2018-2022 79ae158ceab8e4e083c8cfcacde56669407a94a2 · Access: restricted
Reported results
DD -0.036 [-0.116, 0.044], p 0.389, Holm p 0.778.
Uncertainty and replication
95% paired item bootstrap interval; Holm correction over three benchmarks.
Limitations
Retrospective design: the GSM8K-PL token cap was raised and capped answers were regenerated at 1280 tokens without a recorded amendment, the latter after an interim grading. MLX bf16 greedy decoding, compared within one numerical stack; GSM8K-PL is machine-translated.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Values from results/analysis.json, per_benchmark_DD.pes.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

PES examination questions

250 multiple-choice questions of the Polish State Specialization Examination 2018 to 2022, allocated proportionally over 57 medical specializations.

Key pes_questions · version a164a223-a32f-4fdf-a694-c1d5cda9b77d

Estimate with interval

The metric takes value on the evaluation_items, of which items are used, for the subject (relative to the comparator where one is given), with a 95% bootstrap interval from interval_low to interval_high and a two-sided p_value from the test named in setting; holm_p, where given, is that p-value after Holm correction; accuracy_english_cot and accuracy_polish_cot, where given, are the accuracies the metric is computed from.

Key estimate_with_interval · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

English chain-of-thought gain difference

English chain-of-thought gain of the subject model minus that of the comparator model, on the same questions.

Key english_cot_gain_difference · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

related

e9fd6b0d-e665-41be-a6dc-290547a16735

The benchmark is built on questions of the same examination.