Experiment · E4
Where do Claude Haiku, Sonnet, Opus and Fable stand on ZendoBench 1.0.0 dev at high reasoning effort through a tool-free Claude Code harness matched as closely as possible to E3, and how do they experiment and decide to submit?
Descriptive inference-only baselines: claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; high effort; all 460 dev games per model, episode seeds from the manifest and no model seed; provider-hosted inference orchestrated on Apple silicon. Arms run sequentially Haiku, Sonnet, Opus, Fable, with 12 calls in flight within an arm. No training or sealed evaluation.
Prerequisites and protocol
- Access needs
- Claude subscription access to all four specified models and macOS for the isolation sandbox. Login and raw run files stay private. Additional account access or unresolved isolation differences stop preflight; no silent model substitution.
- Suggested protocol
- Match E3: unpatched ZendoBench 1.0.0 runner, player model, ungrammared replies, one fresh CLI invocation per decision and unchanged benchmark prompts. Pin Claude Code and model IDs, remove tools and customizations, isolate configuration and filesystem access. Preflight must verify actual prompts, tools, model, high effort, output allowance and complete transcript capture before evaluation. Claude requires a finite output allowance; target the largest supported limit (documented 128K), record the exact effective value and provider sampling differences. Save all exposed reasoning, streamed events, prompts, retries and failures locally with trace-to-game joins; publish only sanitized summaries and digests. Score every run with --verify and compare the same dev games against E3. Register attempts before launch; commit DESIGN.md before measurements.
Accepted plan
Plan accepted by @stw2
E4 remaining runs: continue Haiku alongside Fable, retaining completed Sonnet and Opus- requires ZendoBench 1.0.0 dev manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · Download
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 22
15 preflight failed · 1 started · 3 succeeded · 3 failed · 1 outcome unknown
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
A17 claude-haiku-5-5-claude-high: six-game smoke preflight, then all 460 dev games at high effort; same attempt across unfinished-game continuation parts.
Preflight failedplanned attempt · stw2/zendo-lab @ 55723010f403 · under an accepted plan
A18 claude-sonnet-5-5-claude-high: six-game smoke preflight, then all 460 dev games at high effort; same attempt across unfinished-game continuation parts.
Preflight failedplanned attempt · stw2/zendo-lab @ 55723010f403 · under an accepted plan
A19 claude-opus-5-5-claude-high: six-game smoke preflight, then all 460 dev games at high effort; same attempt across unfinished-game continuation parts.
Preflight failedplanned attempt · stw2/zendo-lab @ 55723010f403 · under an accepted plan
A21 Haiku high-effort baseline after A17 temporary-rate-limit preflight failure. No prior model/game outcome. Connection check, six-game smoke, then 460 dev games; unchanged settings.
Preflight failedrerun attempt · stw2/zendo-lab @ c60e22fcc4db · under an accepted plan
rerun attempt · stw2/zendo-lab @ c578f63052a0 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan
A28 claude-sonnet-5-5-claude-high: six-game smoke then full 460-game high-effort baseline using corrected denial audit; launch strictly after preceding arm verifies.
Preflight failedrerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan
A29 claude-opus-5-5-claude-high: six-game smoke then full 460-game high-effort baseline using corrected denial audit; launch strictly after preceding arm verifies.
Preflight failedrerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan
A30 claude-fable-5-1-claude-high: six-game smoke then full 460-game high-effort baseline using corrected denial audit; launch strictly after preceding arm verifies.
Preflight failedrerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan
rerun attempt · stw2/zendo-lab @ c4bf8e221582 · under an accepted plan
A37 claude-fable-5-1-claude-high-v2: six-game smoke, verification and full 460-game dev baseline; starts after Haiku, Sonnet and Opus all verify.
Preflight failedrerun attempt · stw2/zendo-lab @ c4bf8e221582 · under an accepted plan
A38 claude-fable-5-1-claude-high-v2: six-game smoke, score --verify and transcript audit, then full 460-game dev baseline; overlaps the remaining Haiku A31 games. Sonnet A35 and Opus A36 are already complete and retained.
Started · outcome unknownrerun attempt · stw2/zendo-lab @ 552752353a9e · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find Learning to test hypotheses, thread Campaign: can training teach small models to test hypotheses?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →