Immutable accepted plan · prospective
Paired tool-use trajectories and one-shot per-step probes of Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, compared with a step-composition prediction
The design locked on 7 July 2026 at commit c85a1ec, after the task instances were frozen and before any trajectory or probe was generated.
Public source
https://github.com/stw2/tokenizer-science-tax @ 03e30b6cecb9c899914b18931b1f8ee0bdc7d2f5
Reference checked 2026-09-14 13:17 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- The end-to-end gap exceeds the predicted gap; the gap is larger with English instructions; with English instructions the APT4 model reads more than 1.10 times the original's prompt-side tokens and ends more trajectories by overflow or budget exhaustion.
- Protocol
- 01 builds 100 task instances from 20 templates and freezes the manifest. 03 runs every instance with English and with Polish instructions through a ReAct tool loop with a persistent Python sandbox, one model at a time: at most 8 steps, 700 generated tokens per step, a FINAL answer ends a trajectory only in a turn without code. 02 poses every instance one-shot as a per-step probe. 04 grades trajectories with golden-gated graders (numeric tolerance, hidden tests). 05 computes the end-to-end gap pooled, per language and per family with a paired bootstrap, the language interaction, the English token ratio and non-finalization rates, the step-composition prediction and Holm-adjusted p for H2 and H3. 06 writes flat headline values.
- Dataset
- 100 synthetic task instances generated by 20 parametric templates (numpy PCG64, seed 20260707 plus a template hash and the instance index), each with English and hand-translated Polish instructions; no public benchmark items.
- Split
- None: both models run every instance with both instruction languages.
- Access needs
- Both Bielik 11B repositories on Hugging Face are gated; the two MLX bf16 conversions are restricted materials. Task instances rebuild from the public generators.
- Configurations
- Tool loop: at most 8 steps, 700 generated tokens per step, context window 32,768 tokens, observations cut at 2,000 characters, 20 s per step; one-shot probes: 900 generated tokens; greedy decoding; system prompts calib-2.
- Metric
- End-to-end success gap (original minus APT4 model) pooled, per instruction language and per family with 95% paired bootstrap intervals; the language interaction DD; the English prompt-side token ratio and overflow or budget-exhaustion rates; the predicted gap and D = observed minus predicted gap.
- Seeds
- 20260707 (instance generation base seed and bootstrap)
- Interpretation rule
- H1: confirmed if D, the end-to-end gap minus the predicted gap, is above 0 with a 95% interval excluding 0; refuted if the interval's upper bound is below 0.02; partial otherwise. The interval comes from a paired bootstrap over instances (2000 resamples) with the prediction recomputed in each resample. H2: DD above 0 predicted; Holm-adjusted with H3. H3: English prompt-side token ratio above 1.10 and a higher English overflow or budget-exhaustion rate for the APT4 model; Holm-adjusted with H2. Trajectories that end without an accepted FINAL answer count as failures.
- Resources
- Apple silicon, 128 GB unified memory, MLX bf16, one 11B model resident at a time; at most about 4.5 million generated tokens for trajectories and 0.2 million for probes.
- Prior work
- arXiv:2604.10799v1 compares the pair on static benchmarks only.
Selected exact hypotheses and premises
Hypothesis · H27
0ea2c3bc-bc27-4395-9371-ce71dbf2c962H1: the end-to-end gap exceeds the step-composition prediction.
Hypothesis · H28
3b6df128-98bc-4e74-bbdd-169e6cf68f7dH2: the gap is larger with English instructions.
Hypothesis · H29
320b5b37-af54-49ab-b447-6c352730ad53H3: the APT4 model's English token cost and overflow or budget exhaustion.
Premise · P101
64285728-7520-4c5e-9d7a-9dbbe2c874f7The original model's <think> tag on English prompts, counted as it is under matched budgets.