Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · prospective

E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Prospective measurement amendment after A22 stopped smoke: preserve first complete API response, block CLI repairs without resampling, explicitly request summarized thinking. No finished benchmark games yet; stopped smoke verified and archived.

Public source

Plan

Prediction
Descriptive baselines; no directional model-ranking hypothesis.
Protocol
DESIGN.md and dated amendments fix the protocol. Sequential Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1 at high effort. One fixed SDK identity sentence precedes the unchanged benchmark system message; exact user message; all environment/date/budget wrapper text removed. Empty tools, OS sandbox, guarded requests, full private transcripts. Per arm: separate six-game smoke then 460 dev games, 12 concurrent stateless calls. Score --verify and paired E3 comparisons. Return the first complete provider response unchanged, including malformed text, regardless of a blocked CLI continuation. Native continuations never reach the provider. Actual structured tool blocks still stop the run. Explicit thinking configuration is adaptive with display summarized.
Dataset
ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
Split
dev; all 460 games; no train or sealed evaluation
Access needs
Claude subscription with access to the four explicit model IDs; macOS sandbox. Dedicated login verified. Raw trajectories and transcripts remain restricted to the owner.
Configurations
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
Metric
Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
Seeds
Episode seeds from the fixed dev manifest; no model sampling seed.
Interpretation rule
Descriptive comparisons, explicitly carrying fixed-SDK-identity, sampling and finite-cap differences from E3. Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm.
Resources
Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
Prior work
E3 A12-A16 are existing Codex baselines. E4 A17/A21 exact-prompt preflights failed; minimal fixed SDK identity controls succeeded. A22 passed guarded connection but its smoke stopped on blocked CLI continuation, with zero finished games under score --verify. All artifacts retained. Before measuring, offline fixtures verify explicit summaries, first-response preservation and interruption cleanup. No full-arm measurements have started.

Selected exact hypotheses and premises

No hypotheses or premises selected.