Finding

Sign in with GitHub
← Publications

Finding · P99 · Author-curated

On needle retrieval over Polish scientific documents at unconditional accuracy 0.70, the effective context is at least 106,235 characters for Bielik-PL-11B-v3.0-Instruct and 0 characters for Bielik-11B-v3.0-Instruct, whose accuracy is below 0.70 at every length, so the ratio of the two is undefined; in the 875 of 2,000 bootstrap resamples in which it is defined, its 95% interval runs from 1.4998 to 28.1940.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Context comparison with an undefined ratio

metric
Effective context in charactersQuantity compared.
language
PolishLanguage of the documents.
threshold
0.70 accuracyAccuracy threshold of the metric.
subject
Bielik-PL-11B-v3.0-InstructModel in the ratio's numerator.
subject value
106235 charactersSubject's effective context.
subject lower bound
trueWhether subject_value is censored at the largest measured length.
comparator
Bielik-11B-v3.0-InstructModel in the ratio's denominator.
comparator value
0 charactersComparator's effective context.
resamples defined
875 resamplesBootstrap resamples in which the ratio is defined.
interval low
1.4998 ratioLower bound of the 95% bootstrap interval of the ratio over those resamples.
interval high
28.1940 ratioUpper bound of the 95% bootstrap interval of the ratio over those resamples.

Experimental provenance

Method and evaluation protocol
Unconditional accuracy per length over 30 prompts (6 needles at 5 depths), prompts over the token limit scored incorrect; effective context is the largest length with interpolated accuracy of at least 0.70, and 0 when accuracy is below 0.70 at the shortest length; paired item bootstrap of the ratio.
Dataset
Polish Wikipedia science articles and Polish State Specialization Examination questions from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
Unconditional accuracy by length, Bielik-PL-11B-v3.0-Instruct: 3,768: 1.000; 18,839: 0.833; 37,678: 1.000; 56,517: 0.967; 69,328: 1.000; 81,384: 1.000; 86,605: 0.967; 106,235: 1.000. Bielik-11B-v3.0-Instruct: 3,768: 0.667; 18,839: 0.567; 37,678: 0.100; 56,517: 0.000; 69,328: 0.467; 81,384: 0.000 (30 of 30 over the limit); 86,605: 0.000 (30 of 30 over the limit); 106,235: 0.000 (30 of 30 over the limit). Ratio at threshold 0.50: 4.9342; at 0.85: undefined.
Uncertainty and replication
95% paired item bootstrap over 6 items, 2000 resamples, of which 875 give a defined ratio; one-sided p 0.0011 and Holm-adjusted 0.0023, computed over those resamples only.
Limitations
Under the locked rule the point ratio is undefined, so the test gives no verdict; resamples without a defined ratio are left out of the interval and the p value. Bielik-11B-v3.0-Instruct's accuracy on this task is below 0.70 from the shortest length, an instruction-following failure rather than a context limit. Polish documents include examination questions that some answers follow. The two models differ in continued pretraining and post-training as well as in tokenizer.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Amendment 3 of the design, dated 5 July 2026 and written at analysis after this result was seen, adds a fallback floor: the larger of this interval's lower bound and the censored cloze-retrieval ratio (1.3961), that is 1.4998, binned as partial agreement with nearly doubled capacity. That reading is secondary to the locked rule.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Context comparison with an undefined ratio

On the task in the language at the threshold, the subject's value of the metric is subject_value, a lower bound when subject_lower_bound is true, and the comparator's is comparator_value, which is 0, so the ratio of subject_value to comparator_value is undefined; interval_low and interval_high bound the 95% bootstrap interval of the ratio over the resamples_defined resamples in which it is defined.

Key context_comparison_undefined_ratio · version ccaa3bc9-94b6-4197-9fd6-993bda035bea

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Effective context in characters

Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.

Key effective_context_characters · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Needle retrieval in scientific documents

Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Exact references

related

95ab68f2-8a5f-4597-aed0-30be3374eae2

Why Bielik-11B-v3.0-Instruct's effective context on this task is 0.

related

96bbca50-8834-4704-b367-656432b83da8

The same models and language on the secondary task.

related

b0861e86-822f-4f13-92c2-0e8f7704010d

Character ceilings of the same models on the same documents.

related

9ac91b29-87d3-48c1-a32f-95da879874b3

The locked bins for the claim read this ratio, so the claim gets no verdict. Secondary note: Amendment 3, dated 5 July 2026 and written after this result was seen, reads a floor of 1.4998, in the 1.3 to 1.8 bin for partial agreement.