Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · retrospective

Accuracy against characters of context for the Bielik 11B pair on English and Polish scientific documents

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 3, dated 5 July 2026, at analysis after the Polish needle-retrieval ratio was seen to be undefined and before findings were written: a fallback floor for the adjudication. The order rests on file modification times. Arms 05 and 06 were added in the port to reproduce the reasoning-tag counts and the flat values.

Public source

Plan

Prediction
English: the ratio of Bielik-11B-v3.0-Instruct's effective context in characters to Bielik-PL-11B-v3.0-Instruct's exceeds 1, about 1.5 to 1.6. Polish: the reverse ratio exceeds 1, about 1.55. At the largest English length both models fit, Bielik-PL-11B-v3.0-Instruct is less accurate.
Protocol
01 measures each model's character ceiling at the 32,704-token prompt limit by binary search on the real assembly, emits the character grid and assembles every prompt deterministically; 02 runs smoke gates on 14 prompts twice, then greedy generation for every prompt that fits, one model at a time, with a generation budget of 1024 tokens capped per prompt at the 32,768-token window; prompts over the limit score incorrect without generation; 03 grades answers against golden cases by the last code or number candidate in the output, or the single candidate of the first non-empty line; 04 computes unconditional and within-window accuracy curves with Wilson intervals, effective context in characters at thresholds 0.50, 0.70 and 0.85, capacity ratios with a paired item bootstrap, the paired within-window contrast, Holm adjustment over the three tests and the adjudication of the nearly doubled capacity claim.
Dataset
Needle retrieval: 6 needles (3 codes, 3 numbers) at 5 depths per language and length, in English LaTeX method sections and arXiv abstracts and in Polish Wikipedia science articles and PES examination questions. Cloze retrieval: 10 items per language at 3 depths among headed distractor documents. No-context controls for both tasks.
Split
None: every assembled prompt is run once per model; cloze items that either model answers without documents are dropped from the cloze analysis.
Access needs
Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face; two corpora and the prompts and generations built from them are restricted materials; generation needs Apple silicon with MLX.
Configurations
Needle-retrieval lengths: 0.05, 0.25, 0.50, 0.75, 0.92 and 1.08 of the smaller character ceiling and 0.75 and 0.92 of the larger; cloze-retrieval lengths: 0.25, 0.60 and 0.92 of the smaller and 0.92 of the larger; depths 0.10 to 0.90; lengths within 5 percent merged; generation budget 1024 tokens; sizing rule: over 18 projected hours 4 needles and 6 cloze items, under 9 hours 8 needles.
Metric
Effective context in characters at unconditional accuracy 0.70 per model, language and task; capacity ratio with a 95 percent paired item bootstrap interval; paired within-window accuracy difference at the largest English length both models fit.
Seeds
20260704 (assembly, needle values and bootstrap)
Interpretation rule
Holm-adjusted one-sided tests over three hypotheses: English ratio above 1, Polish ratio above 1, within-window difference below 0. The Polish ratio at threshold 0.70 adjudicates the claim that capacity nearly doubles: at least 1.8 confirms, 1.3 to 1.8 partial, 1.1 to 1.3 partial trending to refute, below 1.1 refutes. Gates that must pass on smoke before the full run: BOS parity, byte-identical determinism on 6 prompts per model, token-count agreement, all golden cases, tiny-context accuracy at least 0.90 per model and language, no-context needle accuracy at most 1 in 6. When the Polish needle-retrieval ratio at 0.70 is undefined, the adjudication uses a floor, the larger of the Polish cloze-retrieval ratio (its lower bound when censored) and the lower bound of the needle-retrieval bootstrap interval, binned as above, with confirms barred unless the floor is at least 1.8.
Resources
Apple silicon, 128 GB unified memory; MLX bf16; about 8 to 16 hours of generation.
Prior work
Section 7 of arXiv:2604.10799v1 derives nearly doubled Polish context capacity from preamble fertility; corpus-level fertility of the same documents was measured before this design.

Selected exact hypotheses and premises

Hypothesis · H10

54782e3d-eed7-4fa0-9b9d-932294a9001a

The English capacity ratio.

Hypothesis · H11

2720c2cd-8b97-4995-9676-87c4a0ae42aa

The Polish capacity ratio.

Hypothesis · H12

c3768eb0-ac40-4420-8fd4-1416d55bb2cb

The within-window contrast.

Premise · P18

9ac91b29-87d3-48c1-a32f-95da879874b3

The claim tested as a capability.