Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · prospective

E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Prospective minimal-wrapper amendment before benchmark games: retain the fixed SDK identity sentence required by the working Claude subscription preflight; preserve benchmark messages and empty tools. Capture cleanup also verified.

Public source

Plan

Prediction
Descriptive baselines; no directional model-ranking hypothesis.
Protocol
DESIGN.md and dated amendments fix the protocol. Sequential Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1 at high effort. One fixed SDK identity sentence precedes the unchanged benchmark system message; exact user message; all environment/date/budget wrapper text removed. Empty tools, OS sandbox, guarded requests, full private transcripts. Per arm: separate six-game smoke then 460 dev games, 12 concurrent stateless calls. Score --verify and paired E3 comparisons.
Dataset
ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
Split
dev; all 460 games; no train or sealed evaluation
Access needs
Claude subscription with access to the four explicit model IDs; macOS sandbox. Dedicated login verified. Raw trajectories and transcripts remain restricted to the owner.
Configurations
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.
Metric
Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
Seeds
Episode seeds from the fixed dev manifest; no model sampling seed.
Interpretation rule
Descriptive comparisons, explicitly carrying fixed-SDK-identity, sampling and finite-cap differences from E3. Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm.
Resources
Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
Prior work
E3 A12-A16 are existing Codex baselines. E4 offline fixtures pass all four IDs and interruption cleanup. A17/A21 exact-prompt preflights failed with HTTP 429; A21 standard-wrapper and minimal-SDK-identity diagnostic controls answered OK. No benchmark games or results. All diagnostics retained privately with hashes.

Selected exact hypotheses and premises

No hypotheses or premises selected.