Finding

Sign in with GitHub
← Publications

Finding · P102 · Author-curated

Asked the 12 English and Polish needle questions without documents, Bielik-PL-11B-v3.0-Instruct answers 0 correctly and Bielik-11B-v3.0-Instruct 0.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Answers correct

task
Needle retrieval in scientific documentsTask whose questions are asked.
setting
needle questions of both languages asked without documentsConditions of the questions.
total
12 questionsQuestions asked of each model.
subject
Bielik-PL-11B-v3.0-InstructModel compared.
subject correct
0 questionsSubject's correct answers.
comparator
Bielik-11B-v3.0-InstructModel compared with.
comparator correct
0 questionsComparator's correct answers.

Experimental provenance

Method and evaluation protocol
The 6 needle questions per language asked with no documents and graded like the main prompts.
Dataset
Needle questions without documents; the gold codes and numbers are generated.Version: unspecified · Access: public
Reported results
Bielik-PL-11B-v3.0-Instruct 0/12; Bielik-11B-v3.0-Instruct 0/12. Cloze items dropped because a model answered them without documents: English 1, Polish 3.
Uncertainty and replication
Exact counts.
Evidence references
analysis.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/c60abf85776b3d991bdbba0e47569c6d944056e7/experiments/E04-effective-context/results/analysis.jsongraded_pl.jsonl · restrictedNot redistributed: ask the Room owner, or rebuild with scripts/03_grade.py at the pinned revisionsgraded_orig.jsonl · restrictedNot redistributed: ask the Room owner, or rebuild with scripts/03_grade.py at the pinned revisionsmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/c60abf85776b3d991bdbba0e47569c6d944056e7/experiments/E04-effective-context/results/metrics.json
Limitations
6 questions per language; the gate allowed at most 1 correct in 6.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Answers correct

Of total questions of the task asked in the setting, the subject answers subject_correct correctly and the comparator comparator_correct.

Key answers_correct · version 0db4abc4-0c31-4a70-a9b9-3cfc201244a1

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Needle retrieval in scientific documents

Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a