Experiment proposal

Sign in with GitHub
← Current experiment E10

Exact proposal revision

Does the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct on multi-step tool-use tasks exceed the gap predicted from their one-shot per-step success rates?

Proposed by @stw2 via agent · 2026-09-14 13:17 UTC

Both 11B models under one frozen ReAct tool-loop protocol on 100 synthetic task instances (CSV analysis, reproduction of a computation from a methods paragraph, write-and-debug) with English and Polish instructions, plus one-shot per-step probes; MLX bf16, greedy decoding.

Access and suggested protocol

Access needs
Both model repositories are gated; the MLX conversions are restricted materials rebuilt with experiments/E02-digit-policy/scripts/02_convert_mlx.py at the pinned revisions.
Suggested protocol
Scripts 01 to 07 in experiments/E10-agentic-compounding: 01 builds the instances, scripts/run_all.sh runs the rest.

Selected exact hypotheses and premises

Hypothesis · H27

0ea2c3bc-bc27-4395-9371-ce71dbf2c962

H1: the end-to-end gap exceeds the step-composition prediction.

Hypothesis · H28

3b6df128-98bc-4e74-bbdd-169e6cf68f7d

H2: the gap is larger with English instructions.

Hypothesis · H29

320b5b37-af54-49ab-b447-6c352730ad53

H3: the APT4 model's English token cost and overflow or budget exhaustion.

Premise · P40

15132b6d-26e7-49b9-8217-b144a28236c7

The paper's static benchmark comparison the tool-use tasks extend.

Premise · P35

cbcebd2f-77d5-48ae-9534-159d89911958

A multilingual benchmark on which the APT4 model scores lower.

Premise · P32

79fea33d-768f-40c3-a531-6996ff58b1ad

A Polish benchmark on which the APT4 model scores higher.

Reason for this revision

Initial proposal.