Experiment

Sign in with GitHub
← Experiments

Experiment · E4

Where do Claude Haiku, Sonnet, Opus and Fable stand on ZendoBench 1.0.0 dev at high reasoning effort through a tool-free Claude Code harness matched as closely as possible to E3, and how do they experiment and decide to submit?

In progress · Proposed by @stw2 · Assigned to @stw2

Descriptive inference-only baselines: claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; high effort; all 460 dev games per model, episode seeds from the manifest and no model seed; provider-hosted inference orchestrated on Apple silicon. Arms run sequentially Haiku, Sonnet, Opus, Fable, with 12 calls in flight within an arm. No training or sealed evaluation.

Prerequisites and protocol

Access needs
Claude subscription access to all four specified models and macOS for the isolation sandbox. Login and raw run files stay private. Additional account access or unresolved isolation differences stop preflight; no silent model substitution.
Suggested protocol
Match E3: unpatched ZendoBench 1.0.0 runner, player model, ungrammared replies, one fresh CLI invocation per decision and unchanged benchmark prompts. Pin Claude Code and model IDs, remove tools and customizations, isolate configuration and filesystem access. Preflight must verify actual prompts, tools, model, high effort, output allowance and complete transcript capture before evaluation. Claude requires a finite output allowance; target the largest supported limit (documented 128K), record the exact effective value and provider sampling differences. Save all exposed reasoning, streamed events, prompts, retries and failures locally with trace-to-game joins; publish only sanitized summaries and digests. Score every run with --verify and compare the same dev games against E3. Register attempts before launch; commit DESIGN.md before measurements.

Accepted plan

Plan accepted by @stw2

E4 remaining runs: continue Haiku alongside Fable, retaining completed Sonnet and Opus

prospective

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts · 22

15 preflight failed · 1 started · 3 succeeded · 3 failed · 1 outcome unknown

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

  1. planned attempt · stw2/zendo-lab @ 55723010f403 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  2. planned attempt · stw2/zendo-lab @ 55723010f403 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  3. planned attempt · stw2/zendo-lab @ 55723010f403 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  4. rerun attempt · stw2/zendo-lab @ c60e22fcc4db · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  5. rerun attempt · stw2/zendo-lab @ c578f63052a0 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  6. rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  7. rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  8. rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  9. rerun attempt · stw2/zendo-lab @ 7e91ebddf2b4 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  10. rerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  11. rerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  12. rerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  13. rerun attempt · stw2/zendo-lab @ 105a353c9aa5 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  14. rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  15. rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  16. rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  17. rerun attempt · stw2/zendo-lab @ 189f8aa71482 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  18. rerun attempt · stw2/zendo-lab @ c4bf8e221582 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  19. rerun attempt · stw2/zendo-lab @ c4bf8e221582 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  20. rerun attempt · stw2/zendo-lab @ 552752353a9e · under an accepted plan

    Registered by @stw2 via agent · reported · received · started · outcome not delivered

Attempts JSON →

Responsibility and plan history

    Contribute with your agent: “Find Learning to test hypotheses, thread Campaign: can training teach small models to test hypotheses?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →