Finding

Sign in with GitHub
← Publications

Finding · P85 · Author-curated

On 800 Polish STEM questions, Bielik-PL-11B-v3.0-Instruct generates 337.6 tokens per answer on average with Polish chain-of-thought and 234.9 with English chain-of-thought.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Mean generated tokensQuantity compared.
benchmark
Polish STEM question setItems the values are measured on.
scope
Bielik-PL-11B-v3.0-InstructModel generating the answers.
subject
Polish chain-of-thoughtPrompting condition.
subject value
337.6 tokensMean generated tokens under the subject condition.
comparator
English chain-of-thoughtPrompting condition compared with.
comparator value
234.9 tokensMean generated tokens under the comparator condition.
setting
few-shot; greedy decoding; MLX bf16; capped answers regenerated with a 1280-token capDecoding.

Experimental provenance

Method and evaluation protocol
Mean of the generated-token counts of the answers, per model and condition, over all 800 questions.
Dataset
800 Polish STEM questions: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM and 250 PES examination questions.Version: Eurolingua/gsm8kx 363d87617cd417b731dc860194a6935fd5ff49ae; openai/gsm8k 740312add88f781978c0658806c59bc2815b9866; amu-cai/llmzszl-dataset 9c11d52a415147f1e2c44038a9d4d4155f064e1a; speakleash/PES-2018-2022 79ae158ceab8e4e083c8cfcacde56669407a94a2 · Access: restricted
Reported results
337.6 Polish, 234.9 English; difference 102.7.
Uncertainty and replication
Means without intervals.
Limitations
Token counts include the final-answer line and depend on the cap. The pre-registered comparison of the token surplus with the gain across benchmark and model combinations was not run.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Values from results/analysis.json, trace_fertility and pl_trace_token_surplus.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Mean generated tokens

Mean number of tokens a model generates per answer, counted by its own tokenizer.

Key mean_generated_tokens · version 5907d582-517e-4a59-bdc6-c58cf7a7fe7d

English chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in English; the question is in Polish.

Key english_chain_of_thought · version 5907d582-517e-4a59-bdc6-c58cf7a7fe7d

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Polish STEM question set

800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.

Key polish_stem_questions · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Polish chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

Key polish_chain_of_thought · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Exact references

related

8368afe3-5846-42c2-9352-ffa0ff2db353

Tokenizer fertility on Polish text, which the token counts of Polish reasoning reflect.