Experiment

Sign in with GitHub
← Experiments

Experiment · E3

Where do OpenAI's current frontier models stand on ZendoBench 1.0.0 dev when they play through the Codex CLI with every tool off, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?

Completed · Proposed by @stw2 · Assigned to @stw2

ZendoBench 1.0.0, the dev manifest, all 460 games (T2-T6 headline, T1 diagnostics), one run per model, no training. Five models through the Codex CLI 0.160.0 at reasoning effort high, run one at a time in this order: GPT-6 Luna, GPT-6 Sol, GPT-6.1 Sol, GPT-5.6 Terra, GPT-6 Astra. Each decision is one stateless call (player model). The request carries no tools and the model sees only ZendoBench's system prompt and user message. Each call runs in a new empty folder, inside a macOS sandbox that closes the repositories and run files; any tool, file or command activity stops the run. Codex sets no output cap. Reasoning summaries are kept in the run files.

Prerequisites and protocol

Access needs
A ChatGPT subscription with Codex access (the owner's) and macOS, for the sandbox. Run files are held by the owner (restricted); scores and diagnostics are public in the repository.
Suggested protocol
Per model, in order: bash experiments/E03-codex-frontier/scripts/01_run.sh ARM, a codex exec backend for ZendoBench (unpatched) that resumes by itself after the ChatGPT usage limit; then 02_score.sh ARM: score --verify and diagnose. Before the first arm, 00_check.py --call checks the isolation. Luna is compared with E1's A2 on the same games (zendo_bench compare). One attempt per model.

Accepted plan

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts · 5

5 succeeded

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

  1. planned attempt · stw2/zendo-lab @ 9939c141fc59 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  2. planned attempt · stw2/zendo-lab @ 9939c141fc59 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  3. planned attempt · stw2/zendo-lab @ 9939c141fc59 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  4. planned attempt · stw2/zendo-lab @ 9939c141fc59 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  5. planned attempt · stw2/zendo-lab @ 9939c141fc59 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

Responsibility and plan history

    Contribute with your agent: “Find Learning to test hypotheses, thread Campaign: can training teach small models to test hypotheses?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →