Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · prospective

E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Prospective full-measurement source update after reviewed A23 smoke audit false positive; carry forward verified smoke without repeating inference. Correct only denied-installation-probe classification and E4 PID attribution.

Public source

Plan

Prediction
Descriptive baselines; no directional model-ranking hypothesis.
Protocol
DESIGN.md and dated amendments fix the protocol. Sequential Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1 at high effort. One fixed SDK identity sentence precedes the unchanged benchmark system message; exact user message; all environment/date/budget wrapper text removed. Empty tools, OS sandbox, guarded requests, full private transcripts. Per arm: separate six-game smoke then 460 dev games, 12 concurrent stateless calls. Score --verify and paired E3 comparisons. Return the first complete provider response unchanged, including malformed text, regardless of a blocked CLI continuation. Native continuations never reach the provider. Actual structured tool blocks still stop the run. Explicit thinking configuration is adaptive with display summarized. Audit-only amendment: attribute denials to recorded E4 CLI PIDs and classify only the three exact blocked native installation directory reads as expected; no sandbox permission or model-facing change. Reuse A23 verified six-game Haiku smoke after its original exit-5 audit was explained and revised audit passed; start fresh full-run attempts. Other model smokes remain mandatory.
Dataset
ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
Split
dev; all 460 games; no train or sealed evaluation
Access needs
Claude subscription with access to the four explicit model IDs; macOS sandbox. Dedicated login verified. Raw trajectories and transcripts remain restricted to the owner.
Configurations
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
Metric
Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
Seeds
Episode seeds from the fixed dev manifest; no model sampling seed.
Interpretation rule
Descriptive comparisons, explicitly carrying fixed-SDK-identity, sampling and finite-cap differences from E3. Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm.
Resources
Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
Prior work
E3 A12-A16 are existing Codex baselines. E4 A17/A21 exact-prompt preflights failed; minimal fixed SDK identity controls succeeded. A22 passed guarded connection but its smoke stopped on blocked CLI continuation, with zero finished games under score --verify. All artifacts retained. Before measuring, offline fixtures verify explicit summaries, first-response preservation and interruption cleanup. No full-arm measurements have started. A23 completed six Haiku smoke games and passed score --verify and complete transcript joins; the wrapper audit stopped on three denied native-installation cleanup directory reads. Pinned binary source and E4 PID attribution explain these; corrected audit passes. Original files and exit 5 are retained. No full baseline has started; A24-A26 were never launched. Model adapter source, prompts, effort, settings and sandbox are unchanged.

Selected exact hypotheses and premises

No hypotheses or premises selected.