Exact proposal revision
Does the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct on multi-step tool-use tasks exceed the gap predicted from their one-shot per-step success rates?
Both 11B models under one frozen ReAct tool-loop protocol on 100 synthetic task instances (CSV analysis, reproduction of a computation from a methods paragraph, write-and-debug) with English and Polish instructions, plus one-shot per-step probes; MLX bf16, greedy decoding.
Access and suggested protocol
- Access needs
- Both model repositories are gated; the MLX conversions are restricted materials rebuilt with experiments/E02-digit-policy/scripts/02_convert_mlx.py at the pinned revisions.
- Suggested protocol
- Scripts 01 to 07 in experiments/E10-agentic-compounding: 01 builds the instances, scripts/run_all.sh runs the rest.
Selected exact hypotheses and premises
Hypothesis · H27
0ea2c3bc-bc27-4395-9371-ce71dbf2c962H1: the end-to-end gap exceeds the step-composition prediction.
Hypothesis · H28
3b6df128-98bc-4e74-bbdd-169e6cf68f7dH2: the gap is larger with English instructions.
Hypothesis · H29
320b5b37-af54-49ab-b447-6c352730ad53H3: the APT4 model's English token cost and overflow or budget exhaustion.
Premise · P40
15132b6d-26e7-49b9-8217-b144a28236c7The paper's static benchmark comparison the tool-use tasks extend.
Premise · P35
cbcebd2f-77d5-48ae-9534-159d89911958A multilingual benchmark on which the APT4 model scores lower.
Premise · P32
79fea33d-768f-40c3-a531-6996ff58b1adA Polish benchmark on which the APT4 model scores higher.
Reason for this revision
Initial proposal.