Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · prospective

E4 remaining runs: continue Haiku alongside Fable, retaining completed Sonnet and Opus

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Owner confirmed renewed authentication and instructed Haiku and Fable to run in parallel. Preserve completed games and models; change only the outer schedule.

Public source

Plan

Prediction
Descriptive baselines; no directional model-ranking hypothesis.
Protocol
Owner-requested 2026-10-09 scheduling amendment: resume Haiku A31 only on its remaining 205 dev games, excluding the 255 verified completed games across five immutable parts. Its original six-game smoke remains verified; do not repeat it. Launch Fable concurrently, with its own separate six-game smoke, score --verify and transcript audit before the full 460-game dev measurement. Retain Sonnet A35 and Opus A36, both complete and verified on all 460 games, and skip their inference and terminal attempt events entirely. Twelve stateless calls per active arm, at most 24 total. Verify all measurement/scoring source bytes against their registered commits; prompts, model IDs, high effort, adaptive summarized thinking, provider-default sampling, max_tokens 128000, fixed SDK identity, empty tools, OS sandbox, bounded retries and transcript capture stay unchanged. Return the first complete API response unchanged, block CLI continuations, and preserve capped or malformed output without resampling. Separate per-arm artifacts and serialize aggregate scoring writes. The owner renewed the dedicated subscription login after expiration; no API-key fallback. Every full run is scored with --verify and paired against E3. Earlier partial/failed attempts remain separate from v2 headline results.
Dataset
ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
Split
dev; all 460 games; no train or sealed evaluation
Access needs
Renewed isolated Claude subscription login; Fable access and settings still require its smoke check. No API-key fallback; raw transcripts remain owner-restricted.
Configurations
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
Metric
Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
Seeds
Episode seeds from the fixed dev manifest; no model sampling seed.
Interpretation rule
Descriptive cross-provider comparisons with fixed SDK identity, provider-default sampling, finite output cap and serving differences disclosed. Preserve A31 original attempt/plan provenance with this scheduling amendment; unchanged configuration permits continuation without resampling finished games. Sonnet/Opus are retained as completed measurements; Fable is prospective. Scheduling overlap can affect timing and failure rates; no latency or equal-compute claim. Unexplained guard violations stop evaluation. Preserve interrupted games and their transcripts separately from completed outcomes.
Resources
Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestration; 12 calls per active arm, up to two active arms (24 calls). Shared subscription capacity may limit throughput.
Prior work
Sonnet A35 and Opus A36 have completed all 460 dev games, score --verify, transcript audits and paired E3 comparisons, and have succeeded Room receipts. Their results were available before this amendment. Haiku A31 has 255 verified completed games and complete transcript joins across five retained parts, then stopped on repeated revoked-token HTTP 401 responses; native renewal failed. The owner confirmed a new isolated sign-in and explicitly requested Haiku/Fable overlap. The stored credential is now unexpired and the isolated CLI reports subscription authentication. A37 Fable has never launched and made no model calls; it will be superseded by a registration pinned to this amendment. Earlier A27 partial results remain excluded from v2. Seven coordinator tests pass, covering parallel launch, failure gating, continuation without repeated smoke, live-worker adoption and skipping completed arms.

Selected exact hypotheses and premises

No hypotheses or premises selected.