Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · retrospective

Accuracy against characters of context for the Bielik 11B pair on English and Polish scientific documents

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

The design dated 4 July 2026 as locked before generation, without its amendments. It was first committed with the amendments and the results on 5 July 2026, so the plan is retrospective. The versions of scripts 02 to 04 that ran before the amendments were not kept; each enters the repository with the amendment its surviving version implements.

Public source

Plan

Prediction
English: the ratio of Bielik-11B-v3.0-Instruct's effective context in characters to Bielik-PL-11B-v3.0-Instruct's exceeds 1, about 1.5 to 1.6. Polish: the reverse ratio exceeds 1, about 1.55. At the largest English length both models fit, Bielik-PL-11B-v3.0-Instruct is less accurate.
Protocol
01 measures each model's character ceiling at the 32,704-token prompt limit by binary search on the real assembly, emits the character grid and assembles every prompt deterministically; 02 runs smoke gates on 14 prompts twice, then greedy generation for every prompt that fits, one model at a time, with a generation budget of 64 tokens; prompts over the limit score incorrect without generation; 03 grades answers against golden cases by the last code or number candidate in the output; 04 computes unconditional and within-window accuracy curves with Wilson intervals, effective context in characters at thresholds 0.50, 0.70 and 0.85, capacity ratios with a paired item bootstrap, the paired within-window contrast, Holm adjustment over the three tests and the adjudication of the nearly doubled capacity claim.
Dataset
Needle retrieval: 6 needles (3 codes, 3 numbers) at 5 depths per language and length, in English LaTeX method sections and arXiv abstracts and in Polish Wikipedia science articles and PES examination questions. Cloze retrieval: 10 items per language at 3 depths among headed distractor documents. No-context controls for both tasks.
Split
None: every assembled prompt is run once per model; cloze items that either model answers without documents are dropped from the cloze analysis.
Access needs
Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face; two corpora and the prompts and generations built from them are restricted materials; generation needs Apple silicon with MLX.
Configurations
Needle-retrieval lengths: 0.05, 0.25, 0.50, 0.75, 0.92 and 1.08 of the smaller character ceiling and 0.75 and 0.92 of the larger; cloze-retrieval lengths: 0.25, 0.60 and 0.92 of the smaller and 0.92 of the larger; depths 0.10 to 0.90; lengths within 5 percent merged; generation budget 64 tokens; sizing rule: over 18 projected hours 4 needles and 6 cloze items, under 9 hours 8 needles.
Metric
Effective context in characters at unconditional accuracy 0.70 per model, language and task; capacity ratio with a 95 percent paired item bootstrap interval; paired within-window accuracy difference at the largest English length both models fit.
Seeds
20260704 (assembly, needle values and bootstrap)
Interpretation rule
Holm-adjusted one-sided tests over three hypotheses: English ratio above 1, Polish ratio above 1, within-window difference below 0. The Polish ratio at threshold 0.70 adjudicates the claim that capacity nearly doubles: at least 1.8 confirms, 1.3 to 1.8 partial, 1.1 to 1.3 partial trending to refute, below 1.1 refutes. Gates that must pass on smoke before the full run: BOS parity, byte-identical determinism on 6 prompts per model, token-count agreement, all golden cases, tiny-context accuracy at least 0.90 per model and language, no-context needle accuracy at most 1 in 6.
Resources
Apple silicon, 128 GB unified memory; MLX bf16; about 8 to 16 hours of generation.
Prior work
Section 7 of arXiv:2604.10799v1 derives nearly doubled Polish context capacity from preamble fertility; corpus-level fertility of the same documents was measured before this design.

Selected exact hypotheses and premises

Hypothesis · H10

54782e3d-eed7-4fa0-9b9d-932294a9001a

The English capacity ratio.

Hypothesis · H11

2720c2cd-8b97-4995-9676-87c4a0ae42aa

The Polish capacity ratio.

Hypothesis · H12

c3768eb0-ac40-4420-8fd4-1416d55bb2cb

The within-window contrast.

Premise · P18

9ac91b29-87d3-48c1-a32f-95da879874b3

The claim tested as a capability.

Premise · P66

81582930-04ae-4afe-809e-f2d13cec4c18

The no-break-space number format costs the APT4 model arithmetic accuracy; prompts normalise it at assembly.