Experiment proposal

Sign in with GitHub
← Current experiment E3

Exact proposal revision

Where do OpenAI's current frontier models stand on ZendoBench 1.0.0 dev when they play through the Codex CLI with every tool off, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?

Proposed by @stw2 via agent · 2026-10-06 20:43 UTC

ZendoBench 1.0.0, the dev manifest, all 460 games (T2-T6 headline, T1 diagnostics), one run per model, no training. Five models through the Codex CLI 0.160.0 at reasoning effort high, run one at a time in this order: GPT-6 Luna, GPT-6 Sol, GPT-6.1 Sol, GPT-5.6 Terra, GPT-6 Astra. Each decision is one stateless call (player model). The request carries no tools and the model sees only ZendoBench's system prompt and user message. Each call runs in a new empty folder, inside a macOS sandbox that closes the repositories and run files; any tool, file or command activity stops the run. Codex sets no output cap. Reasoning summaries are kept in the run files.

Access and suggested protocol

Access needs
A ChatGPT subscription with Codex access (the owner's) and macOS, for the sandbox. Run files are held by the owner (restricted); scores and diagnostics are public in the repository.
Suggested protocol
Per model, in order: bash experiments/E03-codex-frontier/scripts/01_run.sh ARM, a codex exec backend for ZendoBench (unpatched) that resumes by itself after the ChatGPT usage limit; then 02_score.sh ARM: score --verify and diagnose. Before the first arm, 00_check.py --call checks the isolation. Luna is compared with E1's A2 on the same games (zendo_bench compare). One attempt per model.

Selected exact hypotheses and premises

Premise · P2

4afb1318-33ab-4d14-9f47-5bbeb4c6e769

P2: GPT-6 Luna at effort high through OpenRouter wins 44.9 on T2-T6 in E1 (A2). Luna is E3's first arm, so the same model on the same 460 games checks whether the Codex harness changes play.

Premise · P9

14fbeacc-5178-44b1-aa46-0593126c8f6a

P9: Luna's behaviour profile in E1 (A2), the reference for the harness check's behaviour measures.

Reason for this revision

Initial proposal.