Mean generated tokens
Mean number of tokens a model generates per answer, counted by its own tokenizer.
Key mean_generated_tokens · version 5907d582-517e-4a59-bdc6-c58cf7a7fe7d
Finding
Sign in with GitHubFinding · P85 · Author-curated
Relation: Metric comparison
Values from results/analysis.json, trace_fertility and pl_trace_token_surplus.
Reuse the defining version and key when the meaning fits your assertion.
Mean number of tokens a model generates per answer, counted by its own tokenizer.
Key mean_generated_tokens · version 5907d582-517e-4a59-bdc6-c58cf7a7fe7d
Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in English; the question is in Polish.
Key english_chain_of_thought · version 5907d582-517e-4a59-bdc6-c58cf7a7fe7d
The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.
Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e
800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.
Key polish_stem_questions · version a8a88632-e5bd-42b6-8a04-17b77ce87d13
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04
Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.
Key polish_chain_of_thought · version d0b79462-c553-4a1f-9612-98ca6d4a0645
related
8368afe3-5846-42c2-9352-ffa0ff2db353Tokenizer fertility on Polish text, which the token counts of Polish reasoning reflect.
Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.