Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · prospective

E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Owner-approved prospective Claude baseline protocol; offline fixture and isolation checks completed, no live inference yet.

Public source

Plan

Prediction
Descriptive baselines; no directional model-ranking hypothesis.
Protocol
DESIGN.md fixes the protocol. Sequential arms Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Per arm: separate six-game smoke, then all 460 dev games with 12 concurrent stateless decisions. Unmodified ZendoBench 1.0.0, exact system/user prompts, no tools, dedicated subscription login, OS file sandbox, validate requests before forwarding. Full incremental transcripts include retries and partial output. Score --verify and paired comparisons to E3 A12-A16.
Dataset
ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
Split
dev; all 460 games; no train or sealed evaluation
Access needs
Claude subscription with access to the four explicit model IDs; macOS sandbox. Dedicated login verified. Raw trajectories and transcripts remain restricted to the owner.
Configurations
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1
Metric
Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
Seeds
Episode seeds from the fixed dev manifest; no model sampling seed.
Interpretation rule
Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm.
Resources
Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
Prior work
E3 A12-A16 are existing Codex baselines on these dev games. E4 has only offline mock-request checks, no live model output.

Selected exact hypotheses and premises

No hypotheses or premises selected.