Execution attempt

Sign in with GitHub
← Experiment E3 · Where do OpenAI's current frontier models stand on ZendoBench 1.0.0 dev when they play through the Codex CLI with every tool off, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?

Execution attempt · A12 · planned

E3 arm gpt-6-luna-codex-high: all 460 dev games through the Codex CLI with no tools, run per the accepted plan

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/zendo-lab @ 9939c141fc598a90e9d05b4735ce584dc8277bdd

Reference checked 2026-10-06 20:44 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-luna-codex-high
Working directory
.
Configuration paths
experiments/E03-codex-frontier/scripts/sandbox.sb experiments/E03-codex-frontier/scripts/codex-cli/package-lock.json
Parameters
model=gpt-6-luna; backend=Codex CLI 0.160.0 (codex exec, no tools), 12 calls in flight; hardware=provider; sampling=reasoning effort high, summary detailed (Codex defaults: temperature 1.0, top_p 0.98, verbosity low); output cap=none
Environment
macOS 26.5 on Apple M4 Max 128 GB (harness only); Python 3.13 via uv; ZendoBench 1.0.0; Codex CLI 0.160.0 (codex exec, no tools) under sandbox-exec, a new empty folder per call; ChatGPT sign-in in a Codex home of its own; the model served by OpenAI
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

T1.wins
29
T2.wins
45
T3.wins
113
T4.wins
2
T5.wins
14
T6.wins
8
T1.scored
30
T2.scored
60
T3.scored
150
T4.scored
60
T5.scored
100
T6.scored
60
T1.win_rate
0.9666666666666667
T2.win_rate
0.75
T3.win_rate
0.7533333333333334
T4.win_rate
0.03333333333333333
T5.win_rate
0.14
T6.win_rate
0.13333333333333333
diagnose.eig
0.4150315652558824
diagnose.zero
0.48480565371024736
headline_items
430
headline_t2_t6
0.362
headline_ci_low
0.3198652641018461
diagnose.p_first
0.5750612094661641

and 11 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    E3 arm gpt-6-luna-codex-high: all 460 dev games through the Codex CLI with no tools, run per the accepted plan

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Launched 2026-10-06T20:45:25Z at 9939c14 (clean tree), detached: bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-luna-codex-high. Isolation checks passed before launch (results/isolation-check.txt at 9939c14). The run resumes into ARM.partN.jsonl after a ChatGPT usage-limit stop, as the same attempt.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Exit 0; 430/430 headline games finished, score --verify verified every game. T2-T6 headline 36.2 [32.0, 41.1]; seed-only 4.7; 15.4 experiments per game at 0.42 bits of expected information; median 2 classes alive at the first submission. Isolation flags: none. Run files and call traces restricted (sha256 given); summaries public at d351665. Exit codes by part: [1, 0]. Part 1 stopped (exit 1) on an error reading one call's trace after 4,055 calls; its 11 games in flight were replayed in part 2 as the same attempt (DESIGN.md amendment of 2026-10-06; fixed at 76dfab0 for the later arms). Harness check, paired with E1's A2 on the same 460 games: A2 (Luna through OpenRouter) wins 8.7 points more on T2-T6 [4.4, 13.0], mostly on T5 and T6 (luna-codex-vs-a2.compare.md); under the plan's rule the harness changes play.

    Exit code 0. Duration 9509 s.

    T1.wins=29 · T2.wins=45 · T3.wins=113 · T4.wins=2 · T5.wins=14 · T6.wins=8 · T1.scored=30 · T2.scored=60 · …

    • gpt-6-luna-codex-high.jsonl · experiments/E03-codex-frontier/results/runs/gpt-6-luna-codex-high.jsonl (run file; held by the Room owner; not public; ask the reporter) · restricted · sha256 17fbd1a92220… · 9216646 bytes · M41
    • gpt-6-luna-codex-high.part2.jsonl · experiments/E03-codex-frontier/results/runs/gpt-6-luna-codex-high.part2.jsonl (run file; held by the Room owner; not public; ask the reporter) · restricted · sha256 088c3540fd64… · 9304489 bytes · M42
    • gpt-6-luna-codex-high.traces.jsonl.gz · experiments/E03-codex-frontier/results/runs/traces/gpt-6-luna-codex-high.traces.jsonl.gz (call traces; held by the Room owner; not public; ask the reporter) · restricted · sha256 b0578c6f57cf… · 31811099 bytes · M43
    • gpt-6-luna-codex-high.part2.traces.jsonl.gz · experiments/E03-codex-frontier/results/runs/traces/gpt-6-luna-codex-high.part2.traces.jsonl.gz (call traces; held by the Room owner; not public; ask the reporter) · restricted · sha256 010986519034… · 41045208 bytes · M44
    • gpt-6-luna-codex-high.score.json · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/gpt-6-luna-codex-high.score.json · public · sha256 08fd09c39150… · 31958 bytes · M45
    • gpt-6-luna-codex-high.diagnose.md · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/gpt-6-luna-codex-high.diagnose.md · public · sha256 0d70350ef5bf… · 1625 bytes · M46
    • gpt-6-luna-codex-high.files.json · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/gpt-6-luna-codex-high.files.json · public · sha256 17356405ee4d… · 303 bytes · M47
    • luna-codex-vs-a2.compare.json · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/luna-codex-vs-a2.compare.json · public · sha256 44c15636bab5… · 31019 bytes · M48
    • luna-codex-vs-a2.compare.md · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/luna-codex-vs-a2.compare.md · public · sha256 ec7c9d5ba3a3… · 3983 bytes · M49

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.