Finding

Sign in with GitHub
← Publications

Finding · P93 · Author-curated

Within the 32,704-token prompt limit, needle-retrieval prompts over Polish scientific documents hold at most 115,473 characters of documents for Bielik-PL-11B-v3.0-Instruct and 75,356 for Bielik-11B-v3.0-Instruct, a ratio of 1.5324.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Context comparison

metric
Character ceilingQuantity compared.
task
Needle retrieval in scientific documentsTask whose prompts are assembled.
language
PolishLanguage of the documents.
token limit
32704 tokensPrompt token limit.
subject
Bielik-PL-11B-v3.0-InstructModel with the larger value.
subject value
115473 charactersSubject's character ceiling.
comparator
Bielik-11B-v3.0-InstructModel with the smaller value.
comparator value
75356 charactersComparator's character ceiling.
ratio
1.5324 ratiosubject_value divided by comparator_value.

Experimental provenance

Method and evaluation protocol
Binary search over character budgets on the real prompt assembly (first needle, depth 0.5), counting tokens with each model's tokenizer and chat template against 32,768 minus 64 tokens.
Dataset
Polish Wikipedia science articles and Polish State Specialization Examination questions from frozen corpora, interleaved and packed to character budgets.Version: unspecified · Access: restricted
Reported results
115,473 against 75,356 characters; ratio 1.5324.
Uncertainty and replication
Deterministic token counts on one calibration assembly per language and model.
Limitations
Measured on one calibration assembly; prompts of the same character budget differ in token count, so some prompts below a ceiling overflow and some above it fit.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Context comparison

On the task in the language, the subject's value of the metric is subject_value and the comparator's is comparator_value; ratio is subject_value divided by comparator_value, with its 95% bootstrap interval from interval_low to interval_high where given; threshold is the metric's accuracy threshold and token_limit the prompt token limit, where given; subject_lower_bound, where given, is true when subject_value is the largest measured length rather than a crossing, so that subject_value and ratio are lower bounds.

Key context_comparison · version 538541db-aa28-4110-9de1-f59a7a701952

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Character ceiling

Largest number of characters of documents for which a task's fully assembled prompt, with the needle at depth 0.5, stays within a token limit under a model's tokenizer and chat template, found by binary search on the assembly.

Key character_ceiling · version 538541db-aa28-4110-9de1-f59a7a701952

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Needle retrieval in scientific documents

Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a

Exact references

refines

9ac91b29-87d3-48c1-a32f-95da879874b3

On these Polish documents the token limit admits 1.5324 times as many characters for the APT4 model, not about twice as many.

related

8e524448-ccc5-405f-830d-82c821643387

Token ratio of one of the two Polish document sources.

related

844a2495-9a22-4237-a0fd-4cbffefd78a2

Token ratio of the other Polish document source.