Execution attempt

Sign in with GitHub
← Experiment E4 · Where do Claude Haiku, Sonnet, Opus and Fable stand on ZendoBench 1.0.0 dev at high reasoning effort through a tool-free Claude Code harness matched as closely as possible to E3, and how do they experiment and decide to submit?

Execution attempt · A38 · rerun

claude-fable-5-1-claude-high-v2: six-game smoke, score --verify and transcript audit, then full 460-game dev baseline; overlaps the remaining Haiku A31 games. Sonnet A35 and Opus A36 are already complete and retained.

Started · outcome unknown · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Started with no delivered outcome. The result is unknown until the registrant delivers a success, failure or cancellation event.

Pinned source and configuration

https://github.com/stw2/zendo-lab @ 552752353a9eda80ef94286078af1d2cfa0c4eb0

Reference checked 2026-10-09 07:45 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
bash experiments/E04-claude-baselines/scripts/01_run.sh claude-fable-5-1-claude-high-v2
Working directory
.
Configuration paths
experiments/E04-claude-baselines/DESIGN.md
Parameters
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded", "model": "claude-fable-5-1", "sdk_identity": "You are a Claude agent, built on Anthropic's Claude Agent SDK.", "response_policy": "first complete API response; no CLI repairs", "thinking_display": "summarized"}
Environment
Claude Code 2.1.293; macOS; Apple M4 Max 128 GB local orchestrator; hosted Anthropic inference, provider hardware unknown; isolated Claude subscription login. Owner-requested Haiku/Fable overlap at 12 calls per active arm, at most 24 total. Sonnet and Opus already completed and verified.
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Delivered events

  1. Registered

    #1

    claude-fable-5-1-claude-high-v2: six-game smoke, score --verify and transcript audit, then full 460-game dev baseline; overlaps the remaining Haiku A31 games. Sonnet A35 and Opus A36 are already complete and retained.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Starting separate six-game smoke; full measurement follows only after score --verify and transcript audit pass. Arm claude-fable-5-1-claude-high-v2, model claude-fable-5-1, high effort, adaptive summarized thinking, provider-default sampling, max_tokens 128000, Claude Code 2.1.293 function backend; dev manifest, 460 full-run games, 12 concurrent calls per arm; manifest episode seeds, model unseeded; Apple M4 Max 128 GB local orchestration, hosted inference hardware unknown. Owner-authorized Haiku/Fable overlap; Sonnet and Opus already verified. Fixed SDK identity, disabled tools and full private transcript capture unchanged.

    reported · received · @stw2 via agent · attempt only

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

Only the registrant can report this attempt’s events. Other members can discuss it in the Thread.