GSM8K-PL problems
250 GSM8K test problems in the Polish machine translation of gsm8kx, stratified by the digit length of the gold answer, with openai/gsm8k gold answers.
Key gsm8k_pl_problems · version 080dd7b3-7a21-419e-bf4e-c3770f5cdde6
Finding
Sign in with GitHubFinding · P77 · Author-curated
Relation: Estimate with interval
Values from results/analysis.json, per_benchmark_DD.gsm8k_pl.
Reuse the defining version and key when the meaning fits your assertion.
250 GSM8K test problems in the Polish machine translation of gsm8kx, stratified by the digit length of the gold answer, with openai/gsm8k gold answers.
Key gsm8k_pl_problems · version 080dd7b3-7a21-419e-bf4e-c3770f5cdde6
The metric takes value on the evaluation_items, of which items are used, for the subject (relative to the comparator where one is given), with a 95% bootstrap interval from interval_low to interval_high and a two-sided p_value from the test named in setting; holm_p, where given, is that p-value after Holm correction; accuracy_english_cot and accuracy_polish_cot, where given, are the accuracies the metric is computed from.
Key estimate_with_interval · version a8a88632-e5bd-42b6-8a04-17b77ce87d13
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c
English chain-of-thought gain of the subject model minus that of the comparator model, on the same questions.
Key english_cot_gain_difference · version a8a88632-e5bd-42b6-8a04-17b77ce87d13
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04
Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.