Exact proposal revision
Where do Claude Haiku, Sonnet, Opus and Fable stand on ZendoBench 1.0.0 dev at high reasoning effort through a tool-free Claude Code harness matched as closely as possible to E3, and how do they experiment and decide to submit?
Descriptive inference-only baselines: claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; high effort; all 460 dev games per model, episode seeds from the manifest and no model seed; provider-hosted inference orchestrated on Apple silicon. Arms run sequentially Haiku, Sonnet, Opus, Fable, with 12 calls in flight within an arm. No training or sealed evaluation.
Access and suggested protocol
- Access needs
- Claude subscription access to all four specified models and macOS for the isolation sandbox. Login and raw run files stay private. Additional account access or unresolved isolation differences stop preflight; no silent model substitution.
- Suggested protocol
- Match E3: unpatched ZendoBench 1.0.0 runner, player model, ungrammared replies, one fresh CLI invocation per decision and unchanged benchmark prompts. Pin Claude Code and model IDs, remove tools and customizations, isolate configuration and filesystem access. Preflight must verify actual prompts, tools, model, high effort, output allowance and complete transcript capture before evaluation. Claude requires a finite output allowance; target the largest supported limit (documented 128K), record the exact effective value and provider sampling differences. Save all exposed reasoning, streamed events, prompts, retries and failures locally with trace-to-game joins; publish only sanitized summaries and digests. Score every run with --verify and compare the same dev games against E3. Register attempts before launch; commit DESIGN.md before measurements.
Selected exact hypotheses and premises
No hypotheses or premises selected.
Reason for this revision
Initial proposal.