Finding

Sign in with GitHub
← Publications

Finding · P98 · Author-curated

On document-addressed cloze retrieval over Polish scientific documents, Bielik-11B-v3.0-Instruct's accuracy on prompts within its token limit is at least 0.70 at all 3 measured lengths from 18,839 to 69,328 characters, 0.857 at each.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Accuracy across lengths

metric
Within-window accuracyQuantity measured.
subject
Bielik-11B-v3.0-InstructModel measured.
language
PolishLanguage of the documents.
threshold
0.70 accuracyAccuracy threshold compared with.
lengths below
0 lengthsMeasured lengths with accuracy below the threshold.
lengths measured
3 lengthsMeasured lengths with prompts within the token limit.
shortest length
18839 charactersShortest measured length.
longest length
69328 charactersLongest measured length within the token limit.
lowest accuracy
0.857 accuracyLowest accuracy across the lengths.
highest accuracy
0.857 accuracyHighest accuracy across the lengths.

Experimental provenance

Method and evaluation protocol
Graded generations of prompts within the model's token limit, per length.
Dataset
Polish Wikipedia science articles and Polish State Specialization Examination questions from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
Accuracy within the token limit by length: 18,839: 0.857 over 21 prompts; 45,214: 0.857 over 21 prompts; 69,328: 0.857 over 21 prompts. Bielik-PL-11B-v3.0-Instruct on the same task and language: 18,839: 1.000; 45,214: 0.952; 69,328: 0.905; 106,235: 0.810.
Uncertainty and replication
Point values from one greedy run; 21 prompts per length from the cloze items kept.
Limitations
7 of 10 cloze items remain after the no-context drop rule.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Accuracy across lengths

On the task in the language, the subject's value of the metric is below threshold at lengths_below of lengths_measured lengths from shortest_length to longest_length, and ranges from lowest_accuracy to highest_accuracy across them.

Key accuracy_across_lengths · version 95ab68f2-8a5f-4597-aed0-30be3374eae2

Within-window accuracy

Accuracy of a model on a task over the prompts that fit its token limit.

Key within_window_accuracy · version c3768eb0-ac40-4420-8fd4-1416d55bb2cb

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Document-addressed cloze retrieval

Task in which a real document with a headed code is placed at a set depth among headed distractor documents packed to a character budget, followed by a question that names the document's code, quotes one of its sentences with a number blanked, and asks for the number; items that either model answers without documents are excluded.

Key cloze_retrieval · version 244a9d5c-e698-4591-8bb4-278c88ef9e02

Exact references

related

95ab68f2-8a5f-4597-aed0-30be3374eae2

Same model, language and documents with a trailing needle question.