Finding

Sign in with GitHub
← Publications

Finding · P95 · Author-curated

On document-addressed cloze retrieval over English scientific documents at unconditional accuracy 0.70, the effective context is 95,679 characters for Bielik-11B-v3.0-Instruct and 74,592 characters for Bielik-PL-11B-v3.0-Instruct, a ratio of 1.2827.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Context comparison

metric
Effective context in charactersQuantity compared.
language
EnglishLanguage of the documents.
threshold
0.70 accuracyAccuracy threshold of the metric.
subject
Bielik-11B-v3.0-InstructModel in the ratio's numerator.
subject value
95679.1 charactersSubject's effective context.
subject lower bound
falseWhether subject_value is censored at the largest measured length.
comparator
Bielik-PL-11B-v3.0-InstructModel in the ratio's denominator.
comparator value
74591.8 charactersComparator's effective context.
ratio
1.2827 ratiosubject_value divided by comparator_value.

Experimental provenance

Method and evaluation protocol
Unconditional accuracy per length over 27 prompts (9 items kept of 10 at 3 depths), prompts over the token limit scored incorrect; effective context is the largest length with interpolated accuracy of at least 0.70.
Dataset
English LaTeX method sections and English arXiv abstracts from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
Unconditional accuracy by length, Bielik-11B-v3.0-Instruct: 19,358: 0.889; 46,459: 0.889; 71,237: 0.889; 104,785: 0.630 (3 of 27 over the limit, within-window 0.708). Bielik-PL-11B-v3.0-Instruct: 19,358: 0.889; 46,459: 0.889; 71,237: 0.778 (3 of 27 over the limit, within-window 0.875); 104,785: 0.000 (27 of 27 over the limit). Items dropped because a model answered them without documents: 1.
Uncertainty and replication
No interval: the cloze ratios were not bootstrapped; point estimates from one greedy run.
Limitations
Secondary task, analysed separately and outside the Holm family. Both values are crossings between measured lengths.

Author’s note

The report gives Bielik-PL-11B-v3.0-Instruct's accuracy at 71,237 characters as 0.875, which is its within-window value; the unconditional value is 0.778, corrected in the report on 14 September 2026.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Document-addressed cloze retrieval

Task in which a real document with a headed code is placed at a set depth among headed distractor documents packed to a character budget, followed by a question that names the document's code, quotes one of its sentences with a number blanked, and asks for the number; items that either model answers without documents are excluded.

Key cloze_retrieval · version 244a9d5c-e698-4591-8bb4-278c88ef9e02

Context comparison

On the task in the language, the subject's value of the metric is subject_value and the comparator's is comparator_value; ratio is subject_value divided by comparator_value, with its 95% bootstrap interval from interval_low to interval_high where given; threshold is the metric's accuracy threshold and token_limit the prompt token limit, where given; subject_lower_bound, where given, is true when subject_value is the largest measured length rather than a crossing, so that subject_value and ratio are lower bounds.

Key context_comparison · version 538541db-aa28-4110-9de1-f59a7a701952

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Effective context in characters

Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.

Key effective_context_characters · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Exact references

supports

9d6d1a8a-1772-421f-962b-67f9475cd006

The same ordering of the two models on a second task.