Accepted plan

Sign in with GitHub
← Experiment E3

Immutable accepted plan · prospective

Five OpenAI models through the Codex CLI with every tool off, one at a time (GPT-6 Luna, GPT-6 Sol, GPT-6.1 Sol, GPT-5.6 Terra, GPT-6 Astra), each on the full ZendoBench 1.0.0 dev set (460 games) at reasoning effort high; scored with score --verify and described with diagnose. Luna through Codex is compared with Luna through OpenRouter (E1's A2) on the same games.

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Initial plan.

Public source

Plan

Prediction
Descriptive, no hypothesis. Expected: every arm far above the seed-only baseline (4.7). Luna through Codex within E1 A2's interval (44.9 [40.5, 49.6]); if it is not, the harness changes play.
Protocol
Pre-registered in experiments/E03-codex-frontier/DESIGN.md at this commit. Setup once: scripts/00_setup.py installs Codex CLI 0.160.0 from npm with its lockfile into a state folder outside the repository, with a Codex home and ChatGPT login of its own. It also freezes the server's model catalog (edited per DESIGN.md so that the request carries no tools) and the list of features to disable. Before the first arm: scripts/00_check.py --call (output committed as results/isolation-check.txt). The checks are that the model is shown the user message alone, that the sandbox closes the repositories, run files, ~/.codex and the sealed salt, and that the server's echo of one trivial call shows no tools, effort high and summary detailed. Per arm, in order: SMOKE=1 plays one game per tier as a connection check, not a measurement. Then `bash experiments/E03-codex-frontier/scripts/01_run.sh ARM` writes experiments/E03-codex-frontier/results/runs/ARM.jsonl: all 460 dev items, player model, one stateless codex exec call per decision. ZendoBench's system prompt is passed as Codex's instructions and its user message as the prompt; the last agent message is the reply. A failed call is retried 5 times, then its game is left unfinished and replayed. When the ChatGPT usage limit stops the run, it waits 30 minutes and resumes into ARM.partN.jsonl with --exclude, as the same attempt. An isolation flag stops the run until the cause is understood: a tool, file, command or web item, a file left in the call's folder, or a sandbox denial for the codex process in the system log after each part. Every try of every call is also kept as a call trace beside the run file: Codex's events and the server's messages, including its echo of the request's tools and settings, with no login values. Then `bash experiments/E03-codex-frontier/scripts/02_score.sh ARM`: score --verify and diagnose over the arm's files, and the run files' sha256. One attempt per arm. Run files and call traces are reported as restricted outputs with their sha256 and size; score and diagnose summaries as public outputs at their commit.
Dataset
ZendoBench 1.0.0 dev manifest (zendo_bench/data/bench-v1-dev.json, file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2, manifest hash 6b98d1a5...), from github.com/stw2/zendo-bench tag v1.0.0 (commit 46c192e0ea10a5140a33c1280edb97b0127cc68c).
Split
dev, all 460 items: T1 30, T2 60, T3 150, T4 60, T5 100, T6 60. T2-T6 form the headline; T1 is diagnostics. No train or sealed items.
Access needs
A ChatGPT subscription with Codex access (the owner's); macOS for the sandbox. Run files are held by the owner (restricted); scores and diagnostics are public in the repository.
Configurations
All arms: Codex CLI 0.160.0 (codex exec, binary sha256 112fae7a...), reasoning effort high, reasoning summary detailed, default service tier, 12 calls in flight, player model. The frozen catalog and features are pinned by their sha256 in every run header. Codex's own request defaults, as the server echoed them: temperature 1.0, top_p 0.98, verbosity low, no max_output_tokens. Arms in order: gpt-6-luna-codex-high (gpt-6-luna); gpt-6-sol-codex-high (gpt-6-sol); gpt-6.1-sol-codex-high (gpt-6.1-sol); gpt-5.6-terra-codex-high (gpt-5.6-terra); gpt-6-astra-codex-high (gpt-6-astra).
Metric
Headline: equal-weight mean win rate over T2-T6 with 95% CI (score --verify). Secondary: wins per tier, T1, unscored and unfinished counts, malformed rate, seed-only baseline and w0 bands. From diagnose: experiments per game, bits per experiment, expected information gain, zero-information share, classes alive and p at the first submission, rule parsimony, contradictions, counterexample uptake. Harness check: zendo_bench compare of Luna through Codex with E1's A2. From the run files: calls, calls without an answer, retries, isolation flags.
Seeds
Episode seeds from the dev manifest. The models are not seeded.
Interpretation rule
Descriptive. Arms are compared only where 95% CIs separate. If the paired compare of Luna through Codex with A2 shows a difference outside its interval, the harness changes play, and comparisons of E3's arms with E1's HTTP arms carry that caveat. Any isolation flag stops the run; a confirmed file read by the model voids the arm.
Resources
The owner's ChatGPT subscription through the Codex CLI, and a macOS host for the harness. No GPU.
Prior work
E1 (A2) played GPT-6 Luna at effort high on the same 460 games through OpenRouter. An earlier Codex evaluation (2026-09-23, outside this Room, on a pre-release engine) is not reused.

Selected exact hypotheses and premises

Premise · P2

4afb1318-33ab-4d14-9f47-5bbeb4c6e769

P2: Luna's E1 headline through OpenRouter (A2), the harness check's reference.

Premise · P9

14fbeacc-5178-44b1-aa46-0593126c8f6a

P9: Luna's E1 behaviour profile (A2), the harness check's reference for the diagnose measures.