Immutable accepted plan · retrospective
Accuracy against characters of context for the Bielik 11B pair on English and Polish scientific documents
Amendment 1, dated 4 July 2026, after the smoke runs and before any full-run generation: the generation budget rises from 64 to 1024 tokens, capped per prompt at the 32,768-token window, because Bielik-11B-v3.0-Instruct's reasoning preamble used up 64 tokens before answering. Written into the design file before the results commit; the order rests on file modification times.
Public source
https://github.com/stw2/tokenizer-science-tax @ f5b3bc4d1f83203dd4ae39c2f6fdaf4e7f5ff900
Reference checked 2026-09-14 13:10 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- English: the ratio of Bielik-11B-v3.0-Instruct's effective context in characters to Bielik-PL-11B-v3.0-Instruct's exceeds 1, about 1.5 to 1.6. Polish: the reverse ratio exceeds 1, about 1.55. At the largest English length both models fit, Bielik-PL-11B-v3.0-Instruct is less accurate.
- Protocol
- 01 measures each model's character ceiling at the 32,704-token prompt limit by binary search on the real assembly, emits the character grid and assembles every prompt deterministically; 02 runs smoke gates on 14 prompts twice, then greedy generation for every prompt that fits, one model at a time, with a generation budget of 1024 tokens capped per prompt at the 32,768-token window; prompts over the limit score incorrect without generation; 03 grades answers against golden cases by the last code or number candidate in the output; 04 computes unconditional and within-window accuracy curves with Wilson intervals, effective context in characters at thresholds 0.50, 0.70 and 0.85, capacity ratios with a paired item bootstrap, the paired within-window contrast, Holm adjustment over the three tests and the adjudication of the nearly doubled capacity claim.
- Dataset
- Needle retrieval: 6 needles (3 codes, 3 numbers) at 5 depths per language and length, in English LaTeX method sections and arXiv abstracts and in Polish Wikipedia science articles and PES examination questions. Cloze retrieval: 10 items per language at 3 depths among headed distractor documents. No-context controls for both tasks.
- Split
- None: every assembled prompt is run once per model; cloze items that either model answers without documents are dropped from the cloze analysis.
- Access needs
- Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face; two corpora and the prompts and generations built from them are restricted materials; generation needs Apple silicon with MLX.
- Configurations
- Needle-retrieval lengths: 0.05, 0.25, 0.50, 0.75, 0.92 and 1.08 of the smaller character ceiling and 0.75 and 0.92 of the larger; cloze-retrieval lengths: 0.25, 0.60 and 0.92 of the smaller and 0.92 of the larger; depths 0.10 to 0.90; lengths within 5 percent merged; generation budget 1024 tokens; sizing rule: over 18 projected hours 4 needles and 6 cloze items, under 9 hours 8 needles.
- Metric
- Effective context in characters at unconditional accuracy 0.70 per model, language and task; capacity ratio with a 95 percent paired item bootstrap interval; paired within-window accuracy difference at the largest English length both models fit.
- Seeds
- 20260704 (assembly, needle values and bootstrap)
- Interpretation rule
- Holm-adjusted one-sided tests over three hypotheses: English ratio above 1, Polish ratio above 1, within-window difference below 0. The Polish ratio at threshold 0.70 adjudicates the claim that capacity nearly doubles: at least 1.8 confirms, 1.3 to 1.8 partial, 1.1 to 1.3 partial trending to refute, below 1.1 refutes. Gates that must pass on smoke before the full run: BOS parity, byte-identical determinism on 6 prompts per model, token-count agreement, all golden cases, tiny-context accuracy at least 0.90 per model and language, no-context needle accuracy at most 1 in 6.
- Resources
- Apple silicon, 128 GB unified memory; MLX bf16; about 8 to 16 hours of generation.
- Prior work
- Section 7 of arXiv:2604.10799v1 derives nearly doubled Polish context capacity from preamble fertility; corpus-level fertility of the same documents was measured before this design.
Selected exact hypotheses and premises
Premise · P89
95f8ccf1-de11-482c-99fa-d5240c8f6eddA generation cap cut answers short in the reasoning-language experiment.