Finding

Sign in with GitHub
← Publications

Finding · P91 · Author-curated

On 300 LLMzSzŁ STEM questions with Polish chain-of-thought, after answers that reached 768 tokens are regenerated with a 1280-token cap, 15 answers of Bielik-PL-11B-v3.0-Instruct and 15 of Bielik-11B-v3.0-Instruct reach the cap.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Answers at the token capQuantity compared.
benchmark
LLMzSzŁ STEM questionsItems the values are measured on.
setting
Polish chain-of-thoughtPrompting condition.
scope
answers that reached 768 tokens regenerated with a 1280-token cap; greedy decoding; MLX bf16Decoding.
subject
Bielik-PL-11B-v3.0-InstructModel with the APT4 tokenizer.
subject value
15 answersAnswers of the subject at the cap.
comparator
Bielik-11B-v3.0-InstructModel with the Mistral-derived tokenizer.
comparator value
15 answersAnswers of the comparator at the cap.

Experimental provenance

Method and evaluation protocol
Count answers whose generation stops at the cap in the complete generations after regeneration.
Dataset
300 LLMzSzŁ STEM examination questions.Version: amu-cai/llmzszl-dataset 9c11d52a415147f1e2c44038a9d4d4155f064e1a · Access: restricted
Reported results
15 against 15.
Uncertainty and replication
Exact counts over 300 answers per model.
Limitations
Answers that reached 768 tokens were scored at 1280 tokens while the others kept their first generation.

Author’s note

Values from results/truncation_audit.json, models.<model>.cells["llmzszl_stem|pl"].at_cap_after_extension.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

LLMzSzŁ STEM questions

300 multiple-choice questions in mathematics, physics, science and biology from the LLMzSzŁ test split, excluding vocational examinations and questions that refer to figures or tables.

Key llmzszl_stem_questions · version 4c39157e-9b47-4096-b640-a79bf90a103d

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Answers at the token cap

Number of answers whose generation stops because it reaches the maximum number of generated tokens.

Key answers_at_token_cap · version 95f8ccf1-de11-482c-99fa-d5240c8f6edd

Polish chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

Key polish_chain_of_thought · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

extends

95f8ccf1-de11-482c-99fa-d5240c8f6edd

The same cells after regeneration at the higher cap.