Finding

Sign in with GitHub
← Publications

Finding · P101 · Author-curated

On English needle-retrieval and cloze-retrieval prompts, 110 of 322 generations of Bielik-11B-v3.0-Instruct (0.3416) and 0 of 231 generations of Bielik-PL-11B-v3.0-Instruct contain the <think> tag.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Reasoning tag shareQuantity compared.
language
EnglishLanguage of the prompts.
setting
greedy generation with a budget of up to 1,024 tokens on needle-retrieval and cloze-retrieval prompts within the token limit, without chat-template changesConditions of generation.
subject
Bielik-11B-v3.0-InstructModel compared.
subject value
0.3416 share of generationsSubject's share.
comparator
Bielik-PL-11B-v3.0-InstructModel compared with.
comparator value
0.0000 share of generationsComparator's share.

Experimental provenance

Method and evaluation protocol
Literal match of the tag in each grid generation; no-context controls excluded.
Dataset
Raw generations of both models on English needle-retrieval and cloze-retrieval prompts.Version: unspecified · Access: restricted
Reported results
English: 110/322 against 0/231. Polish: 0/240 for Bielik-11B-v3.0-Instruct and 0/360 for Bielik-PL-11B-v3.0-Instruct. Generations ending at the token budget: 18 of 562 for Bielik-11B-v3.0-Instruct and 9 of 591 for Bielik-PL-11B-v3.0-Instruct.
Uncertainty and replication
Exact counts from one greedy run.
Limitations
The raw generations are not redistributed. The generation budget was raised from 64 to 1,024 tokens after the smoke runs exposed the tag (Amendment 1).

Author’s note

The report gives Bielik-11B-v3.0-Instruct's Polish rate as 0.6%; the same rule gives 0 of 240, corrected in the report on 14 September 2026.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Reasoning tag share

Share of a model's generations whose output contains the literal tag <think>.

Key think_tag_share · version 64285728-7520-4c5e-9d7a-9dbbe2c874f7

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c