Experiment · E10
Does the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct on multi-step tool-use tasks exceed the gap predicted from their one-shot per-step success rates?
Both 11B models under one frozen ReAct tool-loop protocol on 100 synthetic task instances (CSV analysis, reproduction of a computation from a methods paragraph, write-and-debug) with English and Polish instructions, plus one-shot per-step probes; MLX bf16, greedy decoding.
Prerequisites and protocol
- Access needs
- Both model repositories are gated; the MLX conversions are restricted materials rebuilt with experiments/E02-digit-policy/scripts/02_convert_mlx.py at the pinned revisions.
- Suggested protocol
- Scripts 01 to 07 in experiments/E10-agentic-compounding: 01 builds the instances, scripts/run_all.sh runs the rest.
Accepted plan
Plan accepted by @stw2
Paired tool-use trajectories and one-shot per-step probes of Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, compared with a step-composition prediction- requires tasks_manifest.json · configuration · M457 Agentic task manifest · Download
- requires prompts.json · configuration · M458 Tool-loop system prompts calib-2 · Download
- requires data/instances · evaluation data · M459 Agentic task instances · Ask the reporter
- requires pl MLX bf16 conversion · checkpoint · M65 bielik-pl-11b-v3.0-instruct mlx-bf16 · Ask the reporter
- requires orig MLX bf16 conversion · checkpoint · M66 bielik-11b-v3.0-instruct mlx-bf16 · Ask the reporter
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 17
17 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
planned attempt · stw2/tokenizer-science-tax @ 03e30b6cecb9 · under an accepted plan
A93 Runs all 100 instances with English and Polish instructions through the tool loop with Bielik-PL-11B-v3.0-Instruct.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 03e30b6cecb9 · under an accepted plan
A94 Runs all 100 instances with English and Polish instructions through the tool loop with Bielik-11B-v3.0-Instruct.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 03e30b6cecb9 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 03e30b6cecb9 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 03e30b6cecb9 · under an accepted plan
A97 Grades run 1's trajectories with the locked grader, which reads only a written solution.py for write-and-debug tasks.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 03e30b6cecb9 · under an accepted plan
A98 Computes run 1's endpoints from the first grading.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 03e30b6cecb9 · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ df809d2de201 · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ df809d2de201 · under an accepted plan
A101 Writes flat headline values from the analysis, the graded rows and the public trajectories.
Succeededplanned attempt · stw2/tokenizer-science-tax @ df809d2de201 · under an accepted plan
A102 Runs all 100 instances with English and Polish instructions through the tool loop with Bielik-PL-11B-v3.0-Instruct.
Succeededrerun attempt · stw2/tokenizer-science-tax @ d35f2683bfcb · under an accepted plan
A103 Runs all 100 instances with English and Polish instructions through the tool loop with Bielik-11B-v3.0-Instruct.
Succeededrerun attempt · stw2/tokenizer-science-tax @ d35f2683bfcb · under an accepted plan
A104 Poses each instance once, without the tool loop, to Bielik-PL-11B-v3.0-Instruct as a per-step probe in both languages.
Succeededrerun attempt · stw2/tokenizer-science-tax @ d35f2683bfcb · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ d35f2683bfcb · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ d35f2683bfcb · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ d35f2683bfcb · under an accepted plan
A108 Writes flat headline values from the analysis, the graded rows and the public trajectories.
Succeededrerun attempt · stw2/tokenizer-science-tax @ d35f2683bfcb · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Agentic compounding. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →