Immutable accepted plan · retrospective
Paired tool-use trajectories and one-shot per-step probes of Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, compared with a step-composition prediction
Amendment 1 of 8 July 2026, written after run 1's trajectories, probes and first grading and committed with run 1's results: write-and-debug solutions are rebuilt from executed code when no solution.py was written, the partial-run flag counts the manifest's cells, and the analysis names a degenerate prediction. No endpoint changes.
Public source
https://github.com/stw2/tokenizer-science-tax @ df809d2de201ff885d619d31aa541a944f68703c
Reference checked 2026-09-14 13:17 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- Unchanged from the locked design: the end-to-end gap exceeds the predicted gap; the gap is larger with English instructions; the APT4 model's English token cost and overflow or budget exhaustion are higher.
- Protocol
- 01 builds 100 task instances from 20 templates and freezes the manifest. 03 runs every instance with English and with Polish instructions through a ReAct tool loop with a persistent Python sandbox, one model at a time: at most 8 steps, 700 generated tokens per step, a FINAL answer ends a trajectory only in a turn without code. 02 poses every instance one-shot as a per-step probe. 04 grades trajectories with golden-gated graders (numeric tolerance, hidden tests); a write-and-debug solution is rebuilt from the executed code when no solution.py was written. 05 computes the end-to-end gap pooled, per language and per family with a paired bootstrap, the language interaction, the English token ratio and non-finalization rates, the step-composition prediction and Holm-adjusted p for H2 and H3. 06 writes flat headline values.
- Dataset
- 100 synthetic task instances generated by 20 parametric templates (numpy PCG64, seed 20260707 plus a template hash and the instance index), each with English and hand-translated Polish instructions; no public benchmark items.
- Split
- None: both models run every instance with both instruction languages.
- Access needs
- Both Bielik 11B repositories on Hugging Face are gated; the two MLX bf16 conversions are restricted materials. Task instances rebuild from the public generators.
- Configurations
- Tool loop: at most 8 steps, 700 generated tokens per step, context window 32,768 tokens, observations cut at 2,000 characters, 20 s per step; one-shot probes: 900 generated tokens; greedy decoding; system prompts calib-2.
- Metric
- End-to-end success gap (original minus APT4 model) pooled, per instruction language and per family with 95% paired bootstrap intervals; the language interaction DD; the English prompt-side token ratio and overflow or budget-exhaustion rates; the predicted gap and D = observed minus predicted gap.
- Seeds
- 20260707 (instance generation base seed and bootstrap)
- Interpretation rule
- H1: confirmed if D, the end-to-end gap minus the predicted gap, is above 0 with a 95% interval excluding 0; refuted if the interval's upper bound is below 0.02; partial otherwise. The interval comes from a paired bootstrap over instances (2000 resamples) with the prediction recomputed in each resample. H2: DD above 0 predicted; Holm-adjusted with H3. H3: English prompt-side token ratio above 1.10 and a higher English overflow or budget-exhaustion rate for the APT4 model; Holm-adjusted with H2. Trajectories that end without an accepted FINAL answer count as failures.
- Resources
- Apple silicon, 128 GB unified memory, MLX bf16, one 11B model resident at a time; at most about 4.5 million generated tokens for trajectories and 0.2 million for probes.
- Prior work
- arXiv:2604.10799v1 compares the pair on static benchmarks only.
Selected exact hypotheses and premises
Hypothesis · H27
0ea2c3bc-bc27-4395-9371-ce71dbf2c962H1: the end-to-end gap exceeds the step-composition prediction.
Hypothesis · H28
3b6df128-98bc-4e74-bbdd-169e6cf68f7dH2: the gap is larger with English instructions.
Hypothesis · H29
320b5b37-af54-49ab-b447-6c352730ad53H3: the APT4 model's English token cost and overflow or budget exhaustion.
Premise · P101
64285728-7520-4c5e-9d7a-9dbbe2c874f7The original model's <think> tag on English prompts, counted as it is under matched budgets.