PES examination questions
250 multiple-choice questions of the Polish State Specialization Examination 2018 to 2022, allocated proportionally over 57 medical specializations.
Key pes_questions · version a164a223-a32f-4fdf-a694-c1d5cda9b77d
Finding
Sign in with GitHubFinding · P79 · Author-curated
Relation: Estimate with interval
Values from results/analysis.json, per_benchmark_DD.pes.
Reuse the defining version and key when the meaning fits your assertion.
250 multiple-choice questions of the Polish State Specialization Examination 2018 to 2022, allocated proportionally over 57 medical specializations.
Key pes_questions · version a164a223-a32f-4fdf-a694-c1d5cda9b77d
The metric takes value on the evaluation_items, of which items are used, for the subject (relative to the comparator where one is given), with a 95% bootstrap interval from interval_low to interval_high and a two-sided p_value from the test named in setting; holm_p, where given, is that p-value after Holm correction; accuracy_english_cot and accuracy_polish_cot, where given, are the accuracies the metric is computed from.
Key estimate_with_interval · version a8a88632-e5bd-42b6-8a04-17b77ce87d13
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c
English chain-of-thought gain of the subject model minus that of the comparator model, on the same questions.
Key english_cot_gain_difference · version a8a88632-e5bd-42b6-8a04-17b77ce87d13
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04
related
e9fd6b0d-e665-41be-a6dc-290547a16735The benchmark is built on questions of the same examination.
Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.