Accepted plan

Sign in with GitHub
← Experiment E4

Immutable accepted plan · prospective

E4: concurrent Haiku, Sonnet and Opus high-effort baselines, followed by Fable

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

The owner asked to run Haiku, Sonnet and Opus simultaneously. Preserve the running Haiku worker and measurement source; change only coordination and future-arm registration. Fable remains after the first three.

Public source

Plan

Prediction
Descriptive baselines; no directional model-ranking hypothesis.
Protocol
Owner-requested scheduling amendment: Haiku, Sonnet and Opus run concurrently with at most 12 calls per arm, 36 total; Fable waits for all three full runs to verify. A31 Haiku continues in place under its original attempt and unchanged measurement configuration; the outer coordinator adopts its existing child without stopping or rerunning games. Newly launched models retain separate six-game smoke checks, score --verify and transcript audits before their full 460 dev games. Per-arm outputs remain separate; shared aggregate scoring and table writes are serialized. All measurement source files are byte-identical to 189f8aa71482c09ccc7ae8eba765a44952bbeb0c and checked against every attempt registration. DESIGN.md and dated amendments fix the protocol. Each has its own six-game smoke, score --verify and transcript audit, then all 460 dev games with 12 concurrent stateless calls. Fixed SDK identity precedes unchanged benchmark system/user messages; empty tools, OS sandbox and guarded requests. Return first complete API response unchanged, including malformed text and cap stops. Block every CLI continuation/recovery before forwarding. Observed provider streaming/transport failures and HTTP 401/429/server errors qualify for bounded fresh-call wrapper retries; unexplained extra requests still flag. Permit only the isolated runtime config-lock sibling for OAuth renewal. A27 remains a separate partial result; no cross-configuration pooling or carried-forward games in these v2 arms. All raw artifacts remain private and unchanged; paired verified comparisons use E3 baselines.
Dataset
ZendoBench 1.0.0 bench-v1-dev.json, manifest digest sha256:e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2
Split
dev; all 460 games; no train or sealed evaluation
Access needs
Existing isolated Claude subscription login; no API-key fallback. Each newly launched model must pass its smoke check. Raw transcripts remain owner-restricted.
Configurations
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded"}; pinned native Claude Code 2.1.293; model IDs claude-haiku-5-5, claude-sonnet-5-5, claude-opus-5-5, claude-fable-5-1; fixed SDK identity sentence: You are a Claude agent, built on Anthropic's Claude Agent SDK.; thinking.display=summarized; preserve first API response, block CLI continuation
Metric
Equal-weight T2-T6 mean win rate with 95% CI, per-tier results, malformed/unscored/unfinished counts and diagnose measures; paired E3 comparisons.
Seeds
Episode seeds from the fixed dev manifest; no model sampling seed.
Interpretation rule
Descriptive cross-provider comparisons with fixed-SDK-identity, provider-default sampling, finite output cap and serving differences disclosed. Every v2 arm uses one pinned corrected configuration for the full manifest. Preserve and report earlier partial/failed attempts separately; do not pool them into v2 headline results. No outcome-based resampling within an attempt. Unexplained request mismatches and actual tool/access violations stop evaluation; unfinished games retain transcripts and may continue in immutable parts under the same configuration. Scheduling overlap is an operational change; report shared-load timing and failure effects, and make no latency or equal-compute claim. Preserve the A31 original-plan provenance alongside this amendment.
Resources
Hosted Anthropic inference, provider hardware unknown; Apple M4 Max 128 GB local orchestration; 12 concurrent calls per arm, up to three concurrent arms (36 calls). Shared subscription capacity may limit throughput.
Prior work
E3 A12-A16 are completed Codex baselines. Earlier E4 preflight and partial attempts remain preserved. A27 completed 83 verified Haiku games before infrastructure/authentication stops and is excluded from v2. A31 completed its six-game v2 smoke with verification and transcript audit, and its full Haiku run is active with partial outcomes available. The owner now requested concurrent Haiku, Sonnet and Opus. A31 continues under the same attempt and unchanged configuration; this is not a new Haiku run. A32-A34 have made no model calls and are superseded before launch by registrations against this scheduling amendment. Sonnet, Opus and Fable results remain prospective. Three coordinator tests passed and bytewise source checks confirm all measurement and scoring files match the previous pin.

Selected exact hypotheses and premises

No hypotheses or premises selected.