Finding

Sign in with GitHub
← Publications

Finding · P100 · Author-curated

On needle retrieval over English scientific documents at 71,237 characters, the largest measured length both models fit, Bielik-PL-11B-v3.0-Instruct's accuracy is 1.000 and Bielik-11B-v3.0-Instruct's 1.000 on the same 25 prompts, a paired difference of 0.000 (95% bootstrap interval 0.000 to 0.000; Holm-adjusted one-sided p = 1.0).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Within-window accuracyQuantity compared.
language
EnglishLanguage of the documents.
length
71237 charactersLength of the documents in each prompt.
pairs
25 promptsPrompts within both models' token limits.
subject
Bielik-PL-11B-v3.0-InstructModel whose accuracy the difference starts from.
subject value
1.000 accuracySubject's accuracy on the paired prompts.
comparator
Bielik-11B-v3.0-InstructModel whose accuracy is subtracted.
comparator value
1.000 accuracyComparator's accuracy on the paired prompts.
difference
0.000 accuracysubject_value minus comparator_value.
interval low
0.000 accuracyLower bound of the 95% bootstrap interval of the difference.
interval high
0.000 accuracyUpper bound of the 95% bootstrap interval of the difference.

Experimental provenance

Method and evaluation protocol
Correct answers of both models on the prompts at the length that fit both token limits; interval from an item bootstrap; one-sided test of a difference below 0.
Dataset
English LaTeX method sections and English arXiv abstracts from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
25 paired prompts; difference 0.000; one-sided p 1.0; Holm-adjusted 1.0. Within-window accuracy on cloze retrieval, which sits below ceiling: Bielik-11B-v3.0-Instruct 0.708 at 104,785 English characters; Bielik-PL-11B-v3.0-Instruct 0.810 at 106,235 Polish characters.
Uncertainty and replication
95% bootstrap interval over items, 2000 resamples; both models at ceiling, so the interval has zero width.
Evidence references
analysis.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/c60abf85776b3d991bdbba0e47569c6d944056e7/experiments/E04-effective-context/results/analysis.jsongraded_pl.jsonl · restrictedNot redistributed: ask the Room owner, or rebuild with scripts/03_grade.py at the pinned revisionsgraded_orig.jsonl · restrictedNot redistributed: ask the Room owner, or rebuild with scripts/03_grade.py at the pinned revisionsmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/c60abf85776b3d991bdbba0e47569c6d944056e7/experiments/E04-effective-context/results/metrics.json
Limitations
Both models are at ceiling, so the contrast cannot detect within-window degradation; it does not show its absence.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference

On the task in the language at length, over pairs prompts within both models' token limits, the subject's value of the metric is subject_value and the comparator's comparator_value; difference is subject_value minus comparator_value, with its 95% bootstrap interval from interval_low to interval_high.

Key paired_difference · version 61cfeb23-b2da-4d2b-8d81-392b7910d7fc

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Within-window accuracy

Accuracy of a model on a task over the prompts that fit its token limit.

Key within_window_accuracy · version c3768eb0-ac40-4420-8fd4-1416d55bb2cb

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Needle retrieval in scientific documents

Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a