Finding

Sign in with GitHub
← Publications

Finding · P97 · Author-curated

On needle retrieval over Polish scientific documents, Bielik-11B-v3.0-Instruct's accuracy on prompts within its token limit is below 0.70 at all 5 measured lengths from 3,768 to 69,328 characters, ranging from 0.000 to 0.667.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Accuracy across lengths

metric
Within-window accuracyQuantity measured.
subject
Bielik-11B-v3.0-InstructModel measured.
language
PolishLanguage of the documents.
threshold
0.70 accuracyAccuracy threshold compared with.
lengths below
5 lengthsMeasured lengths with accuracy below the threshold.
lengths measured
5 lengthsMeasured lengths with prompts within the token limit.
shortest length
3768 charactersShortest measured length.
longest length
69328 charactersLongest measured length within the token limit.
lowest accuracy
0.000 accuracyLowest accuracy across the lengths.
highest accuracy
0.667 accuracyHighest accuracy across the lengths.

Experimental provenance

Method and evaluation protocol
Graded generations of prompts within the model's token limit, per length.
Dataset
Polish Wikipedia science articles and Polish State Specialization Examination questions from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
Accuracy within the token limit by length: 3,768: 0.667 over 30 prompts; 18,839: 0.567 over 30 prompts; 37,678: 0.100 over 30 prompts; 56,517: 0.000 over 30 prompts; 69,328: 0.467 over 30 prompts. Outcomes over these prompts: correct 54, wrong 35, no answer 46, truncated 15. Bielik-PL-11B-v3.0-Instruct on the same task and language: 3,768: 1.000; 18,839: 0.833; 37,678: 1.000; 56,517: 0.967; 69,328: 1.000; 81,384: 1.000; 86,605: 0.967; 106,235: 1.000.
Uncertainty and replication
Point values from one greedy run; 30 prompts per length from 6 needles.
Limitations
The report and Amendment 3 describe the failed answers as replies offering help without answering the question, read from the raw generations, which are not redistributed; replies by kind were not counted in any stored file. Polish documents include examination questions, which some answers follow. The tiny-context gate passed on two smoke prompts, although accuracy at the shortest length is below its 0.90 halt threshold.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Accuracy across lengths

On the task in the language, the subject's value of the metric is below threshold at lengths_below of lengths_measured lengths from shortest_length to longest_length, and ranges from lowest_accuracy to highest_accuracy across them.

Key accuracy_across_lengths · version 95ab68f2-8a5f-4597-aed0-30be3374eae2

Within-window accuracy

Accuracy of a model on a task over the prompts that fit its token limit.

Key within_window_accuracy · version c3768eb0-ac40-4420-8fd4-1416d55bb2cb

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Needle retrieval in scientific documents

Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a