Finding

Sign in with GitHub
← Publications

Finding · P96 · Author-curated

On document-addressed cloze retrieval over Polish scientific documents at unconditional accuracy 0.70, the effective context is at least 106,235 characters for Bielik-PL-11B-v3.0-Instruct and 76,094 characters for Bielik-11B-v3.0-Instruct, a ratio of at least 1.3961.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Context comparison

metric
Effective context in charactersQuantity compared.
language
PolishLanguage of the documents.
threshold
0.70 accuracyAccuracy threshold of the metric.
subject
Bielik-PL-11B-v3.0-InstructModel in the ratio's numerator.
subject value
106235 charactersSubject's effective context.
subject lower bound
trueWhether subject_value is censored at the largest measured length.
comparator
Bielik-11B-v3.0-InstructModel in the ratio's denominator.
comparator value
76094.3 charactersComparator's effective context.
ratio
1.3961 ratiosubject_value divided by comparator_value.

Experimental provenance

Method and evaluation protocol
Unconditional accuracy per length over 21 prompts (7 items kept of 10 at 3 depths), prompts over the token limit scored incorrect; effective context is the largest length with interpolated accuracy of at least 0.70.
Dataset
Polish Wikipedia science articles and Polish State Specialization Examination questions from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
Unconditional accuracy by length, Bielik-PL-11B-v3.0-Instruct: 18,839: 1.000; 45,214: 0.952; 69,328: 0.905; 106,235: 0.810. Bielik-11B-v3.0-Instruct: 18,839: 0.857; 45,214: 0.857; 69,328: 0.857; 106,235: 0.000 (21 of 21 over the limit). Items dropped because a model answered them without documents: 3.
Uncertainty and replication
No interval: the cloze ratios were not bootstrapped; point estimates from one greedy run.
Limitations
Secondary task, analysed separately and outside the Holm family. The ratio is a lower bound because Bielik-PL-11B-v3.0-Instruct stays at or above 0.70 through the largest length.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Context comparison

On the task in the language, the subject's value of the metric is subject_value and the comparator's is comparator_value; ratio is subject_value divided by comparator_value, with its 95% bootstrap interval from interval_low to interval_high where given; threshold is the metric's accuracy threshold and token_limit the prompt token limit, where given; subject_lower_bound, where given, is true when subject_value is the largest measured length rather than a crossing, so that subject_value and ratio are lower bounds.

Key context_comparison · version 538541db-aa28-4110-9de1-f59a7a701952

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Effective context in characters

Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.

Key effective_context_characters · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Document-addressed cloze retrieval

Task in which a real document with a headed code is placed at a set depth among headed distractor documents packed to a character budget, followed by a question that names the document's code, quotes one of its sentences with a number blanked, and asks for the number; items that either model answers without documents are excluded.

Key cloze_retrieval · version 244a9d5c-e698-4591-8bb4-278c88ef9e02