Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · prospective

E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Pin provider-error retry classification and isolated OAuth lock permission before remaining measurement. Carry A27 verified completed games into an explicit continuation without resampling; partial outcomes disclosed. Await isolated reauthentication before launch.

Public source

Plan

Prediction
Descriptive baselines; no directional model-ranking hypothesis.
Protocol
DESIGN.md and dated amendments fix the protocol. Sequential Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1 at high effort. One fixed SDK identity sentence precedes the unchanged benchmark system message; exact user message; all environment/date/budget wrapper text removed. Empty tools, OS sandbox, guarded requests, full private transcripts. Per arm: separate six-game smoke then 460 dev games, 12 concurrent stateless calls. Score --verify and paired E3 comparisons. Return the first complete provider response unchanged, including malformed text, regardless of a blocked CLI continuation. Native continuations never reach the provider. Actual structured tool blocks still stop the run. Explicit thinking configuration is adaptive with display summarized. Audit-only amendment: attribute denials to recorded E4 CLI PIDs and classify only the three exact blocked native installation directory reads as expected; no sandbox permission or model-facing change. Reuse A23 verified six-game Haiku smoke after its original exit-5 audit was explained and revised audit passed; start fresh full-run attempts. Other model smokes remain mandatory. Provider-error/authentication amendment: keep every second CLI request blocked. An observed provider streaming/transport error or HTTP 401/429/server error becomes a bounded wrapper retry in a fresh stateless call; never resample completed or cap-stopped output. Permit only the isolated config-lock sibling for OAuth renewal. Haiku retains A27 verified completed games as explicit prior inputs and resumes unfinished games; full pooled reporting names both attempt/source histories. No completed game is replayed. Login must be restored before inference.
Dataset
ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
Split
dev; all 460 games; no train or sealed evaluation
Access needs
Restore the isolated Claude subscription login after OAuth expiry; no API-key fallback. Raw transcripts remain owner-restricted. Other model access still requires each smoke check.
Configurations
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
Metric
Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
Seeds
Episode seeds from the fixed dev manifest; no model sampling seed.
Interpretation rule
Descriptive comparisons, explicitly carrying fixed-SDK-identity, sampling and finite-cap differences from E3. Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm. Haiku is pooled across A27 and its registered continuation; declare both sources and attempt provenance, preserve all original completed games, and report unfinished counts. Successful-response semantics, model settings and scoring remain unchanged across the infrastructure correction.
Resources
Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
Prior work
E3 A12-A16 supply existing Codex baselines. E4 earlier preflights and A23 six-game Haiku smoke are recorded with preserved artifacts. A27 completed 83 verified dev games (T1 30, T2 53) under commit 105a353c9aa58ea4dae8f368cbca97e9072acd7d before provider-error recovery and then isolated OAuth expiry/lock-permission stops. Its second part completed zero games. All transcripts finalized and raw parts remain unchanged. These partial outcomes were available before this infrastructure/auth-only correction. A28-A30 never launched. New real-CLI offline fixtures verify blocked recovery, fresh-call failure classification and the exact auth-lock permission. Remaining Haiku measurements and all other model measurements are prospective.

Selected exact hypotheses and premises

No hypotheses or premises selected.