Finding

Sign in with GitHub
← Publications

Finding · P208 · Author-curated

On the 766 of 800 Polish STEM questions usable in the logit lens for both models, with Polish chain-of-thought, Bielik-PL-11B-v3.0-Instruct has answer accuracy 0.7807 and Bielik-11B-v3.0-Instruct 0.7624.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Answer accuracyQuantity compared.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
subject value
0.7807 fractionAccuracy of the subject.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
comparator value
0.7624 fractionAccuracy of the comparator.
setting
Polish chain-of-thoughtPrompting condition of the answers.
scope
766 of 800 Polish STEM questions usable in both models after the lens exclusionsQuestions compared.

Experimental provenance

Method and evaluation protocol
Mean of graded correctness over the questions usable in both models.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
0.7807 against 0.7624.
Uncertainty and replication
Point values without intervals on the lens universe.
Limitations
The universe excludes questions whose answers failed the lens gates in either model.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Answer accuracy

Fraction of questions whose extracted final answer equals the gold answer.

Key answer_accuracy · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Polish chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

Key polish_chain_of_thought · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

related

07d4e6b0-719c-43ec-b75d-ff7bac9cd488

Accuracy of Bielik-PL-11B-v3.0-Instruct on all 800 questions.

related

5696028e-6f29-44d8-a073-8dcd5d4eba07

Accuracy of Bielik-11B-v3.0-Instruct on all 800 questions.