Cited claim

Sign in with GitHub
← Publications

Cited claim · P35 · Author-curated

On INCLUDE-base-44, averaged over 20 European languages, Bielik-PL-11B-v3.0-Instruct scores 53.92 and Bielik-11B-v3.0-Instruct 64.8.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

benchmark
INCLUDE-base-44Benchmark the values are reported on.
scope
average over 20 European languagesPart of the benchmark the value covers.
subject
Bielik-PL-11B-v3.0-InstructModel or tokenizer after the tokenizer change.
subject value
53.92 scoreValue of the metric for the subject.
comparator
Bielik-11B-v3.0-InstructModel or tokenizer before the tokenizer change.
comparator value
64.8 scoreValue of the metric for the comparator.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Section 5.6, p. 13

    For our evaluation, we focus on a subset of 20 European languages from the full benchmark to assess Bielik’s performance across its target linguistic region.
  2. Section 5.6, p. 13

    Bielik-11B-v3.0-Instruct achieves the highest scores among the models listed, with 64.8 average across European languages and 69.0 on Polish-specific tasks.
  3. Section 5.6, p. 13

    The Polish tokenizer variants, Bielik-PL-11B-v3.0-Instruct and Bielik-PL-Minitron-7B-v3.0-Instruct, reach 53.92 and 49.81 on the European-language average and 64.23 and 59.49 on Polish-specific tasks

Author’s note

Comparator values are the leaderboard comparisons from the Bielik 11B v3 technical report (Ociepa et al. [2025a]); subject values are reported in this paper for the checkpoints with the Polish tokenizer.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

INCLUDE-base-44

Multilingual benchmark of four-option multiple-choice questions from academic and professional examinations across 44 languages.

Key include_base_44 · version cbcebd2f-77d5-48ae-9534-159d89911958

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04