Accepted plan

Sign in with GitHub
← Experiment E10

Immutable accepted plan · prospective

Paired tool-use trajectories and one-shot per-step probes of Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, compared with a step-composition prediction

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

The design locked on 7 July 2026 at commit c85a1ec, after the task instances were frozen and before any trajectory or probe was generated.

Public source

Plan

Prediction
The end-to-end gap exceeds the predicted gap; the gap is larger with English instructions; with English instructions the APT4 model reads more than 1.10 times the original's prompt-side tokens and ends more trajectories by overflow or budget exhaustion.
Protocol
01 builds 100 task instances from 20 templates and freezes the manifest. 03 runs every instance with English and with Polish instructions through a ReAct tool loop with a persistent Python sandbox, one model at a time: at most 8 steps, 700 generated tokens per step, a FINAL answer ends a trajectory only in a turn without code. 02 poses every instance one-shot as a per-step probe. 04 grades trajectories with golden-gated graders (numeric tolerance, hidden tests). 05 computes the end-to-end gap pooled, per language and per family with a paired bootstrap, the language interaction, the English token ratio and non-finalization rates, the step-composition prediction and Holm-adjusted p for H2 and H3. 06 writes flat headline values.
Dataset
100 synthetic task instances generated by 20 parametric templates (numpy PCG64, seed 20260707 plus a template hash and the instance index), each with English and hand-translated Polish instructions; no public benchmark items.
Split
None: both models run every instance with both instruction languages.
Access needs
Both Bielik 11B repositories on Hugging Face are gated; the two MLX bf16 conversions are restricted materials. Task instances rebuild from the public generators.
Configurations
Tool loop: at most 8 steps, 700 generated tokens per step, context window 32,768 tokens, observations cut at 2,000 characters, 20 s per step; one-shot probes: 900 generated tokens; greedy decoding; system prompts calib-2.
Metric
End-to-end success gap (original minus APT4 model) pooled, per instruction language and per family with 95% paired bootstrap intervals; the language interaction DD; the English prompt-side token ratio and overflow or budget-exhaustion rates; the predicted gap and D = observed minus predicted gap.
Seeds
20260707 (instance generation base seed and bootstrap)
Interpretation rule
H1: confirmed if D, the end-to-end gap minus the predicted gap, is above 0 with a 95% interval excluding 0; refuted if the interval's upper bound is below 0.02; partial otherwise. The interval comes from a paired bootstrap over instances (2000 resamples) with the prediction recomputed in each resample. H2: DD above 0 predicted; Holm-adjusted with H3. H3: English prompt-side token ratio above 1.10 and a higher English overflow or budget-exhaustion rate for the APT4 model; Holm-adjusted with H2. Trajectories that end without an accepted FINAL answer count as failures.
Resources
Apple silicon, 128 GB unified memory, MLX bf16, one 11B model resident at a time; at most about 4.5 million generated tokens for trajectories and 0.2 million for probes.
Prior work
arXiv:2604.10799v1 compares the pair on static benchmarks only.

Selected exact hypotheses and premises

Hypothesis · H27

0ea2c3bc-bc27-4395-9371-ce71dbf2c962

H1: the end-to-end gap exceeds the step-composition prediction.

Hypothesis · H28

3b6df128-98bc-4e74-bbdd-169e6cf68f7d

H2: the gap is larger with English instructions.

Hypothesis · H29

320b5b37-af54-49ab-b447-6c352730ad53

H3: the APT4 model's English token cost and overflow or budget exhaustion.

Premise · P40

15132b6d-26e7-49b9-8217-b144a28236c7

Static benchmark comparison of the pair.

Premise · P17

eed1981a-09e9-4aa9-bd57-3db32c7ba4a0

English fertility of the two tokenizers.

Premise · P69

d7c031db-103f-410d-b439-1f3ba6cae904

Per-step arithmetic difference of the pair.

Premise · P101

64285728-7520-4c5e-9d7a-9dbbe2c874f7

The original model's <think> tag on English prompts, counted as it is under matched budgets.