Immutable accepted plan · prospective
E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3
Pin provider-error retry classification and isolated OAuth lock permission before remaining measurement. Carry A27 verified completed games into an explicit continuation without resampling; partial outcomes disclosed. Await isolated reauthentication before launch.
Public source
https://github.com/stw2/zendo-lab @ b6fcaacb132eed698fbf27ea243be9f989266ab5
Reference checked 2026-10-08 18:19 UTC. No code was executed or scientific result verified.
- experiments/E04-claude-baselines/DESIGN.md
- experiments/E04-claude-baselines/scripts/claude_backend.py
- experiments/E04-claude-baselines/scripts/01_all.sh
- experiments/E04-claude-baselines/scripts/01_run.sh
- experiments/E04-claude-baselines/scripts/01_play.py
- experiments/E04-claude-baselines/scripts/04_denials.py
- experiments/E04-claude-baselines/scripts/sandbox.sb
- experiments/E04-claude-baselines/scripts/00_check.py
- experiments/E04-claude-baselines/scripts/03_table.py
Plan
- Prediction
- Descriptive baselines; no directional model-ranking hypothesis.
- Protocol
- DESIGN.md and dated amendments fix the protocol. Sequential Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1 at high effort. One fixed SDK identity sentence precedes the unchanged benchmark system message; exact user message; all environment/date/budget wrapper text removed. Empty tools, OS sandbox, guarded requests, full private transcripts. Per arm: separate six-game smoke then 460 dev games, 12 concurrent stateless calls. Score --verify and paired E3 comparisons. Return the first complete provider response unchanged, including malformed text, regardless of a blocked CLI continuation. Native continuations never reach the provider. Actual structured tool blocks still stop the run. Explicit thinking configuration is adaptive with display summarized. Audit-only amendment: attribute denials to recorded E4 CLI PIDs and classify only the three exact blocked native installation directory reads as expected; no sandbox permission or model-facing change. Reuse A23 verified six-game Haiku smoke after its original exit-5 audit was explained and revised audit passed; start fresh full-run attempts. Other model smokes remain mandatory. Provider-error/authentication amendment: keep every second CLI request blocked. An observed provider streaming/transport error or HTTP 401/429/server error becomes a bounded wrapper retry in a fresh stateless call; never resample completed or cap-stopped output. Permit only the isolated config-lock sibling for OAuth renewal. Haiku retains A27 verified completed games as explicit prior inputs and resumes unfinished games; full pooled reporting names both attempt/source histories. No completed game is replayed. Login must be restored before inference.
- Dataset
- ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
- Split
- dev; all 460 games; no train or sealed evaluation
- Access needs
- Restore the isolated Claude subscription login after OAuth expiry; no API-key fallback. Raw transcripts remain owner-restricted. Other model access still requires each smoke check.
- Configurations
- {"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
- Metric
- Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
- Seeds
- Episode seeds from the fixed dev manifest; no model sampling seed.
- Interpretation rule
- Descriptive comparisons, explicitly carrying fixed-SDK-identity, sampling and finite-cap differences from E3. Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm. Haiku is pooled across A27 and its registered continuation; declare both sources and attempt provenance, preserve all original completed games, and report unfinished counts. Successful-response semantics, model settings and scoring remain unchanged across the infrastructure correction.
- Resources
- Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
- Prior work
- E3 A12-A16 supply existing Codex baselines. E4 earlier preflights and A23 six-game Haiku smoke are recorded with preserved artifacts. A27 completed 83 verified dev games (T1 30, T2 53) under commit 105a353c9aa58ea4dae8f368cbca97e9072acd7d before provider-error recovery and then isolated OAuth expiry/lock-permission stops. Its second part completed zero games. All transcripts finalized and raw parts remain unchanged. These partial outcomes were available before this infrastructure/auth-only correction. A28-A30 never launched. New real-CLI offline fixtures verify blocked recovery, fresh-call failure classification and the exact auth-lock permission. Remaining Haiku measurements and all other model measurements are prospective.
Selected exact hypotheses and premises
No hypotheses or premises selected.