Immutable accepted plan · prospective
E4 remaining runs: continue Haiku alongside Fable, retaining completed Sonnet and Opus
Owner confirmed renewed authentication and instructed Haiku and Fable to run in parallel. Preserve completed games and models; change only the outer schedule.
Public source
https://github.com/stw2/zendo-lab @ 552752353a9eda80ef94286078af1d2cfa0c4eb0
Reference checked 2026-10-09 07:45 UTC. No code was executed or scientific result verified.
- experiments/E04-claude-baselines/DESIGN.md
- experiments/E04-claude-baselines/scripts/claude_backend.py
- experiments/E04-claude-baselines/scripts/01_all.sh
- experiments/E04-claude-baselines/scripts/01_run.sh
- experiments/E04-claude-baselines/scripts/01_play.py
- experiments/E04-claude-baselines/scripts/04_denials.py
- experiments/E04-claude-baselines/scripts/sandbox.sb
- experiments/E04-claude-baselines/scripts/00_check.py
- experiments/E04-claude-baselines/scripts/03_table.py
- experiments/E04-claude-baselines/scripts/07_parallel.py
Plan
- Prediction
- Descriptive baselines; no directional model-ranking hypothesis.
- Protocol
- Owner-requested 2026-10-09 scheduling amendment: resume Haiku A31 only on its remaining 205 dev games, excluding the 255 verified completed games across five immutable parts. Its original six-game smoke remains verified; do not repeat it. Launch Fable concurrently, with its own separate six-game smoke, score --verify and transcript audit before the full 460-game dev measurement. Retain Sonnet A35 and Opus A36, both complete and verified on all 460 games, and skip their inference and terminal attempt events entirely. Twelve stateless calls per active arm, at most 24 total. Verify all measurement/scoring source bytes against their registered commits; prompts, model IDs, high effort, adaptive summarized thinking, provider-default sampling, max_tokens 128000, fixed SDK identity, empty tools, OS sandbox, bounded retries and transcript capture stay unchanged. Return the first complete API response unchanged, block CLI continuations, and preserve capped or malformed output without resampling. Separate per-arm artifacts and serialize aggregate scoring writes. The owner renewed the dedicated subscription login after expiration; no API-key fallback. Every full run is scored with --verify and paired against E3. Earlier partial/failed attempts remain separate from v2 headline results.
- Dataset
- ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
- Split
- dev; all 460 games; no train or sealed evaluation
- Access needs
- Renewed isolated Claude subscription login; Fable access and settings still require its smoke check. No API-key fallback; raw transcripts remain owner-restricted.
- Configurations
- {"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
- Metric
- Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
- Seeds
- Episode seeds from the fixed dev manifest; no model sampling seed.
- Interpretation rule
- Descriptive cross-provider comparisons with fixed SDK identity, provider-default sampling, finite output cap and serving differences disclosed. Preserve A31 original attempt/plan provenance with this scheduling amendment; unchanged configuration permits continuation without resampling finished games. Sonnet/Opus are retained as completed measurements; Fable is prospective. Scheduling overlap can affect timing and failure rates; no latency or equal-compute claim. Unexplained guard violations stop evaluation. Preserve interrupted games and their transcripts separately from completed outcomes.
- Resources
- Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestration; 12 calls per active arm, up to two active arms (24 calls). Shared subscription capacity may limit throughput.
- Prior work
- Sonnet A35 and Opus A36 have completed all 460 dev games, score --verify, transcript audits and paired E3 comparisons, and have succeeded Room receipts. Their results were available before this amendment. Haiku A31 has 255 verified completed games and complete transcript joins across five retained parts, then stopped on repeated revoked-token HTTP 401 responses; native renewal failed. The owner confirmed a new isolated sign-in and explicitly requested Haiku/Fable overlap. The stored credential is now unexpired and the isolated CLI reports subscription authentication. A37 Fable has never launched and made no model calls; it will be superseded by a registration pinned to this amendment. Earlier A27 partial results remain excluded from v2. Seven coordinator tests pass, covering parallel launch, failure gating, continuation without repeated smoke, live-worker adoption and skipping completed arms.
Selected exact hypotheses and premises
No hypotheses or premises selected.