Execution attempt

Sign in with GitHub
← Experiment E4 · Where do Claude Haiku, Sonnet, Opus and Fable stand on ZendoBench 1.0.0 dev at high reasoning effort through a tool-free Claude Code harness matched as closely as possible to E3, and how do they experiment and decide to submit?

Execution attempt · A25 · rerun

claude-opus-5-5-claude-high: queued in strict approved order after preceding arm verifies. Fresh six-game smoke then 460 dev games; first API response, no tools or repairs, summarized thinking, fixed SDK identity.

Preflight failed · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/zendo-lab @ 7e91ebddf2b4535046ddcff4c1eb373ca7138a00

Reference checked 2026-10-08 10:22 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
bash experiments/E04-claude-baselines/scripts/01_run.sh claude-opus-5-5-claude-high
Working directory
.
Configuration paths
experiments/E04-claude-baselines/DESIGN.md
Parameters
{"reasoning_effort": "high", "thinking": "adaptive", "max_tokens": 128000, "sampling": "provider defaults", "manifest": "dev", "games": 460, "batch": 12, "player": "model", "backend": "function", "seed": "manifest episode seeds; model unseeded", "model": "claude-opus-5-5", "sdk_identity": "You are a Claude agent, built on Anthropic's Claude Agent SDK.", "response_policy": "first complete API response; no CLI repairs", "thinking_display": "summarized"}
Environment
Claude Code 2.1.293; macOS; Apple M4 Max 128 GB local orchestrator; hosted Anthropic inference, provider hardware unknown; isolated Claude subscription login.
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Delivered events

  1. Registered

    #1

    claude-opus-5-5-claude-high: queued in strict approved order after preceding arm verifies. Fresh six-game smoke then 460 dev games; first API response, no tools or repairs, summarized thinking, fixed SDK identity.

    reported · received · @stw2 via agent · posted to the Thread

  2. Preflight failed before launch

    #2

    A25 was never launched; superseded before inference to pin the corrected sandbox-denial audit. No model call or result.

    Error: superseded_source_before_launch

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.