Immutable accepted plan · prospective
E4 corrected v2 harness: fresh Haiku, Sonnet, Opus and Fable high-effort baselines matched to E3
Before corrected-source inference, native benchmark checks reject mixing harness configurations. Replace the unexecuted carry-forward plan with fresh v2 arm identities, each with its own smoke and complete run; preserve A27 separately. Model IDs/settings unchanged; isolated login required before launch.
Public source
https://github.com/stw2/zendo-lab @ 189f8aa71482c09ccc7ae8eba765a44952bbeb0c
Reference checked 2026-10-08 18:23 UTC. No code was executed or scientific result verified.
- experiments/E04-claude-baselines/DESIGN.md
- experiments/E04-claude-baselines/scripts/claude_backend.py
- experiments/E04-claude-baselines/scripts/01_all.sh
- experiments/E04-claude-baselines/scripts/01_run.sh
- experiments/E04-claude-baselines/scripts/01_play.py
- experiments/E04-claude-baselines/scripts/04_denials.py
- experiments/E04-claude-baselines/scripts/sandbox.sb
- experiments/E04-claude-baselines/scripts/00_check.py
- experiments/E04-claude-baselines/scripts/03_table.py
Plan
- Prediction
- Descriptive baselines; no directional model-ranking hypothesis.
- Protocol
- DESIGN.md and dated amendments fix the protocol. Fresh corrected arms ending in -claude-high-v2 run strictly Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1. Each has its own six-game smoke, score --verify and transcript audit, then all 460 dev games with 12 concurrent stateless calls. Fixed SDK identity precedes unchanged benchmark system/user messages; empty tools, OS sandbox and guarded requests. Return first complete API response unchanged, including malformed text and cap stops. Block every CLI continuation/recovery before forwarding. Observed provider streaming/transport failures and HTTP 401/429/server errors qualify for bounded fresh-call wrapper retries; unexplained extra requests still flag. Permit only the isolated runtime config-lock sibling for OAuth renewal. Restore isolated login before launch. A27 remains a separate partial result; no cross-configuration pooling or carried-forward games in these v2 arms. All raw artifacts remain private and unchanged; paired verified comparisons use E3 baselines.
- Dataset
- ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
- Split
- dev; all 460 games; no train or sealed evaluation
- Access needs
- Restore the isolated Claude subscription login after OAuth expiry; no API-key fallback. Raw transcripts remain owner-restricted. Other model access still requires each smoke check.
- Configurations
- {"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
- Metric
- Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
- Seeds
- Episode seeds from the fixed dev manifest; no model sampling seed.
- Interpretation rule
- Descriptive cross-provider comparisons with fixed-SDK-identity, provider-default sampling, finite output cap and serving differences disclosed. Every v2 arm uses one pinned corrected configuration for the full manifest. Preserve and report earlier partial/failed attempts separately; do not pool them into v2 headline results. No outcome-based resampling within an attempt. Unexplained request mismatches and actual tool/access violations stop evaluation; unfinished games retain transcripts and may continue in immutable parts under the same configuration.
- Resources
- Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
- Prior work
- E3 A12-A16 are existing Codex baselines. E4 earlier preflights and A23 smoke are preserved. A27 completed 83 verified Haiku games under source 105a353c9aa58ea4dae8f368cbca97e9072acd7d before infrastructure/authentication stops; its second part completed none. These partial outcomes were known before the correction. The initial carry-forward plan was not executed: the native ZendoBench configuration identity includes backend source and sandbox configuration and refuses pooling. A27 remains separate; v2 arms start fresh complete runs under one corrected configuration. A28-A30 were never launched. Thirty targeted tests and real-CLI offline fixtures validate protocol enforcement, blocked provider-error recovery and the isolated lock permission; no corrected-source model inference has run.
Selected exact hypotheses and premises
No hypotheses or premises selected.