Immutable accepted plan · prospective
E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3
Prospective amendment after A17 connection preflight: distinguish temporary provider throttling from actual subscription quota. No benchmark game or model answer was produced; request settings unchanged.
Public source
https://github.com/stw2/zendo-lab @ c60e22fcc4dbda7a32d7d7b3ef1f2cfc079619d9
Reference checked 2026-10-08 09:45 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- Descriptive baselines; no directional model-ranking hypothesis.
- Protocol
- DESIGN.md fixes the protocol. Sequential arms Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Per arm: separate six-game smoke, then all 460 dev games with 12 concurrent stateless decisions. Unmodified ZendoBench 1.0.0, exact system/user prompts, no tools, dedicated subscription login, OS file sandbox, validate requests before forwarding. Full incremental transcripts include retries and partial output. Score --verify and paired comparisons to E3 A12-A16.
- Dataset
- ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
- Split
- dev; all 460 games; no train or sealed evaluation
- Access needs
- Claude subscription with access to the four explicit model IDs; macOS sandbox. Dedicated login verified. Raw trajectories and transcripts remain restricted to the owner.
- Configurations
- {"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1
- Metric
- Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
- Seeds
- Episode seeds from the fixed dev manifest; no model sampling seed.
- Interpretation rule
- Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm.
- Resources
- Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
- Prior work
- Offline fixtures passed for all four models. A17 connection preflight returned HTTP 429 without a model answer; rate-limit classification fixed and full trace retained. A18-A20 were never launched and superseded before any call. No E4 benchmark results have been observed.
Selected exact hypotheses and premises
No hypotheses or premises selected.