Finding

Sign in with GitHub
← Publications

Finding · P94 · Author-curated

On needle retrieval over English scientific documents at unconditional accuracy 0.70, the effective context is at least 104,785 characters for Bielik-11B-v3.0-Instruct and 73,219 characters for Bielik-PL-11B-v3.0-Instruct, a ratio of at least 1.4311 (95% bootstrap interval 1.2851 to 1.5883; Holm-adjusted one-sided p = 0.0015).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Context comparison

metric
Effective context in charactersQuantity compared.
language
EnglishLanguage of the documents.
threshold
0.70 accuracyAccuracy threshold of the metric.
subject
Bielik-11B-v3.0-InstructModel in the ratio's numerator.
subject value
104785 charactersSubject's effective context.
subject lower bound
trueWhether subject_value is censored at the largest measured length.
comparator
Bielik-PL-11B-v3.0-InstructModel in the ratio's denominator.
comparator value
73219.4 charactersComparator's effective context.
ratio
1.4311 ratiosubject_value divided by comparator_value.
interval low
1.2851 ratioLower bound of the 95% bootstrap interval of the ratio.
interval high
1.5883 ratioUpper bound of the 95% bootstrap interval of the ratio.

Experimental provenance

Method and evaluation protocol
Unconditional accuracy per length over 30 prompts (6 needles at 5 depths), prompts over the token limit scored incorrect; effective context is the largest length with interpolated accuracy of at least 0.70; interval from a paired item bootstrap.
Dataset
English LaTeX method sections and English arXiv abstracts from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
Unconditional accuracy by length, Bielik-11B-v3.0-Instruct: 3,872: 1.000; 19,358: 1.000; 38,716: 0.933; 58,074: 1.000; 71,237: 1.000; 83,627: 1.000; 104,785: 0.800 (5 of 30 over the limit, within-window 0.960). Bielik-PL-11B-v3.0-Instruct: 3,872: 1.000; 19,358: 1.000; 38,716: 1.000; 58,074: 1.000; 71,237: 0.833 (5 of 30 over the limit, within-window 1.000); 83,627: 0.000 (30 of 30 over the limit); 104,785: 0.000 (30 of 30 over the limit). Ratio at threshold 0.50: 1.3753; at 0.85: 1.4230; excluding truncated generations: 1.4311.
Uncertainty and replication
95% paired item bootstrap over 6 items, 2000 resamples; one-sided p 0.0005 before Holm adjustment over three tests.
Limitations
The ratio is a lower bound because Bielik-11B-v3.0-Instruct stays at or above 0.70 through the largest length. The two models differ in continued pretraining and post-training as well as in tokenizer. One greedy run; accuracy per length rests on 6 needles.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Context comparison

On the task in the language, the subject's value of the metric is subject_value and the comparator's is comparator_value; ratio is subject_value divided by comparator_value, with its 95% bootstrap interval from interval_low to interval_high where given; threshold is the metric's accuracy threshold and token_limit the prompt token limit, where given; subject_lower_bound, where given, is true when subject_value is the largest measured length rather than a crossing, so that subject_value and ratio are lower bounds.

Key context_comparison · version 538541db-aa28-4110-9de1-f59a7a701952

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Effective context in characters

Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.

Key effective_context_characters · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Needle retrieval in scientific documents

Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Exact references

related

538541db-aa28-4110-9de1-f59a7a701952

Character ceilings of the same models on the same documents.