Finding

Sign in with GitHub
← Publications

Finding · P92 · Author-curated

Within the 32,704-token prompt limit, needle-retrieval prompts over English scientific documents hold at most 113,897 characters of documents for Bielik-11B-v3.0-Instruct and 77,432 for Bielik-PL-11B-v3.0-Instruct, a ratio of 1.4709.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Context comparison

metric
Character ceilingQuantity compared.
task
Needle retrieval in scientific documentsTask whose prompts are assembled.
language
EnglishLanguage of the documents.
token limit
32704 tokensPrompt token limit.
subject
Bielik-11B-v3.0-InstructModel with the larger value.
subject value
113897 charactersSubject's character ceiling.
comparator
Bielik-PL-11B-v3.0-InstructModel with the smaller value.
comparator value
77432 charactersComparator's character ceiling.
ratio
1.4709 ratiosubject_value divided by comparator_value.

Experimental provenance

Method and evaluation protocol
Binary search over character budgets on the real prompt assembly (first needle, depth 0.5), counting tokens with each model's tokenizer and chat template against 32,768 minus 64 tokens.
Dataset
English LaTeX method sections and English arXiv abstracts from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
113,897 against 77,432 characters; ratio 1.4709.
Uncertainty and replication
Deterministic token counts on one calibration assembly per language and model.
Limitations
Measured on one calibration assembly; prompts of the same character budget differ in token count, so some prompts below a ceiling overflow and some above it fit.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Character ceiling

Largest number of characters of documents for which a task's fully assembled prompt, with the needle at depth 0.5, stays within a token limit under a model's tokenizer and chat template, found by binary search on the assembly.

Key character_ceiling · version 538541db-aa28-4110-9de1-f59a7a701952

Context comparison

On the task in the language, the subject's value of the metric is subject_value and the comparator's is comparator_value; ratio is subject_value divided by comparator_value, with its 95% bootstrap interval from interval_low to interval_high where given; threshold is the metric's accuracy threshold and token_limit the prompt token limit, where given; subject_lower_bound, where given, is true when subject_value is the largest measured length rather than a crossing, so that subject_value and ratio are lower bounds.

Key context_comparison · version 538541db-aa28-4110-9de1-f59a7a701952

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Needle retrieval in scientific documents

Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Exact references

related

683761d9-7e85-4d6e-9a95-bd924d442e75

Token ratio of one of the two English document sources.

related

490c5f75-43e6-4d68-ba35-6a0097992182

Token ratio of the other English document source.