Cited claim

Sign in with GitHub
← Publications

Cited claim · P37 · Author-curated

On Belebele, averaged over 28 European language variants, Bielik-PL-11B-v3.0-Instruct scores 77.41 and Bielik-11B-v3.0-Instruct 82.98.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

benchmark
BelebeleBenchmark the values are reported on.
scope
average over 28 European language variantsPart of the benchmark the value covers.
subject
Bielik-PL-11B-v3.0-InstructModel or tokenizer after the tokenizer change.
subject value
77.41 scoreValue of the metric for the subject.
comparator
Bielik-11B-v3.0-InstructModel or tokenizer before the tokenizer change.
comparator value
82.98 scoreValue of the metric for the comparator.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Section 5.7, p. 14

    For our evaluation, we assess performance across 28 European language variants to evaluate Bielik’s reading comprehension capabilities across its target linguistic region.
  2. Section 5.7, p. 14

    Bielik-11B-v3.0-Instruct achieves 82.98 average across European languages, representing a substantial improvement over the previous version Bielik-11B-v2.6-Instruct (68.67).
  3. Section 5.7, p. 14

    The Polish tokenizer variants, Bielik-PL-11B-v3.0-Instruct and Bielik-PL-Minitron-7B-v3.0-Instruct, reach 77.41 and 74.23 on the European-language average and 81.22 and 77.44 on Polish-specific tasks, respectively

Author’s note

Comparator values are the leaderboard comparisons from the Bielik 11B v3 technical report (Ociepa et al. [2025a]); subject values are reported in this paper for the checkpoints with the Polish tokenizer.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Belebele

Multilingual multiple-choice reading comprehension benchmark on FLORES-200 passages.

Key belebele · version e0184bf2-dccb-469b-85be-cefff666bd50

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04