Immutable accepted plan · prospective
E4: Haiku, Sonnet, Opus, Fable high-effort tool-free Claude baselines, matched to E3
Prospective full-measurement source update after reviewed A23 smoke audit false positive; carry forward verified smoke without repeating inference. Correct only denied-installation-probe classification and E4 PID attribution.
Public source
https://github.com/stw2/zendo-lab @ 105a353c9aa58ea4dae8f368cbca97e9072acd7d
Reference checked 2026-10-08 13:38 UTC. No code was executed or scientific result verified.
- experiments/E04-claude-baselines/DESIGN.md
- experiments/E04-claude-baselines/scripts/claude_backend.py
- experiments/E04-claude-baselines/scripts/01_all.sh
- experiments/E04-claude-baselines/scripts/01_run.sh
- experiments/E04-claude-baselines/scripts/01_play.py
- experiments/E04-claude-baselines/scripts/04_denials.py
Plan
- Prediction
- Descriptive baselines; no directional model-ranking hypothesis.
- Protocol
- DESIGN.md and dated amendments fix the protocol. Sequential Haiku 5.5, Sonnet 5.5, Opus 5.5, Fable 5.1 at high effort. One fixed SDK identity sentence precedes the unchanged benchmark system message; exact user message; all environment/date/budget wrapper text removed. Empty tools, OS sandbox, guarded requests, full private transcripts. Per arm: separate six-game smoke then 460 dev games, 12 concurrent stateless calls. Score --verify and paired E3 comparisons. Return the first complete provider response unchanged, including malformed text, regardless of a blocked CLI continuation. Native continuations never reach the provider. Actual structured tool blocks still stop the run. Explicit thinking configuration is adaptive with display summarized. Audit-only amendment: attribute denials to recorded E4 CLI PIDs and classify only the three exact blocked native installation directory reads as expected; no sandbox permission or model-facing change. Reuse A23 verified six-game Haiku smoke after its original exit-5 audit was explained and revised audit passed; start fresh full-run attempts. Other model smokes remain mandatory.
- Dataset
- ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
- Split
- dev; all 460 games; no train or sealed evaluation
- Access needs
- Claude subscription with access to the four explicit model IDs; macOS sandbox. Dedicated login verified. Raw trajectories and transcripts remain restricted to the owner.
- Configurations
- {"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
- Metric
- Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
- Seeds
- Episode seeds from the fixed dev manifest; no model sampling seed.
- Interpretation rule
- Descriptive comparisons, explicitly carrying fixed-SDK-identity, sampling and finite-cap differences from E3. Descriptive paired comparisons only. Report sampling and finite-output-cap differences from E3. No resampling model errors. Request mismatch or incomplete transcript stops the run; data-access violation voids affected arm.
- Resources
- Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestrator; 12 calls concurrently within one arm.
- Prior work
- E3 A12-A16 are existing Codex baselines. E4 A17/A21 exact-prompt preflights failed; minimal fixed SDK identity controls succeeded. A22 passed guarded connection but its smoke stopped on blocked CLI continuation, with zero finished games under score --verify. All artifacts retained. Before measuring, offline fixtures verify explicit summaries, first-response preservation and interruption cleanup. No full-arm measurements have started. A23 completed six Haiku smoke games and passed score --verify and complete transcript joins; the wrapper audit stopped on three denied native-installation cleanup directory reads. Pinned binary source and E4 PID attribution explain these; corrected audit passes. Original files and exit 5 are retained. No full baseline has started; A24-A26 were never launched. Model adapter source, prompts, effort, settings and sandbox are unchanged.
Selected exact hypotheses and premises
No hypotheses or premises selected.