Execution attempt · A12 · planned
E3 arm gpt-6-luna-codex-high: all 460 dev games through the Codex CLI with no tools, run per the accepted plan
Pinned source and configuration
https://github.com/stw2/zendo-lab @ 9939c141fc598a90e9d05b4735ce584dc8277bdd
Reference checked 2026-10-06 20:44 UTC. Later commits, branches or plan changes do not retarget this attempt.
- experiments/E03-codex-frontier/DESIGN.md
- experiments/E03-codex-frontier/scripts/codex_backend.py
- experiments/E03-codex-frontier/scripts/sandbox.sb
- experiments/E03-codex-frontier/scripts/00_setup.py
- experiments/E03-codex-frontier/scripts/00_check.py
- experiments/E03-codex-frontier/scripts/01_play.py
- experiments/E03-codex-frontier/scripts/01_run.sh
- experiments/E03-codex-frontier/scripts/02_score.sh
- experiments/E03-codex-frontier/scripts/04_denials.py
- experiments/E03-codex-frontier/scripts/codex-cli/package.json
- experiments/E03-codex-frontier/scripts/codex-cli/package-lock.json
- experiments/E03-codex-frontier/results/isolation-check.txt
- pyproject.toml
- uv.lock
- Command
- bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-luna-codex-high
- Working directory
- .
- Configuration paths
- experiments/E03-codex-frontier/scripts/sandbox.sb experiments/E03-codex-frontier/scripts/codex-cli/package-lock.json
- Parameters
- model=gpt-6-luna; backend=Codex CLI 0.160.0 (codex exec, no tools), 12 calls in flight; hardware=provider; sampling=reasoning effort high, summary detailed (Codex defaults: temperature 1.0, top_p 0.98, verbosity low); output cap=none
- Environment
- macOS 26.5 on Apple M4 Max 128 GB (harness only); Python 3.13 via uv; ZendoBench 1.0.0; Codex CLI 0.160.0 (codex exec, no tools) under sandbox-exec, a new empty folder per call; ChatGPT sign-in in a Codex home of its own; the model served by OpenAI
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · dataset · Download
Latest reported metrics
Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.
- T1.wins
- 29
- T2.wins
- 45
- T3.wins
- 113
- T4.wins
- 2
- T5.wins
- 14
- T6.wins
- 8
- T1.scored
- 30
- T2.scored
- 60
- T3.scored
- 150
- T4.scored
- 60
- T5.scored
- 100
- T6.scored
- 60
- T1.win_rate
- 0.9666666666666667
- T2.win_rate
- 0.75
- T3.win_rate
- 0.7533333333333334
- T4.win_rate
- 0.03333333333333333
- T5.win_rate
- 0.14
- T6.win_rate
- 0.13333333333333333
- diagnose.eig
- 0.4150315652558824
- diagnose.zero
- 0.48480565371024736
- headline_items
- 430
- headline_t2_t6
- 0.362
- headline_ci_low
- 0.3198652641018461
- diagnose.p_first
- 0.5750612094661641
and 11 more in the attempt JSON.
Delivered events
Registered
#1E3 arm gpt-6-luna-codex-high: all 460 dev games through the Codex CLI with no tools, run per the accepted plan
Started
#2Launched 2026-10-06T20:45:25Z at 9939c14 (clean tree), detached: bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-luna-codex-high. Isolation checks passed before launch (results/isolation-check.txt at 9939c14). The run resumes into ARM.partN.jsonl after a ChatGPT usage-limit stop, as the same attempt.
Succeeded
#3Exit 0; 430/430 headline games finished, score --verify verified every game. T2-T6 headline 36.2 [32.0, 41.1]; seed-only 4.7; 15.4 experiments per game at 0.42 bits of expected information; median 2 classes alive at the first submission. Isolation flags: none. Run files and call traces restricted (sha256 given); summaries public at d351665. Exit codes by part: [1, 0]. Part 1 stopped (exit 1) on an error reading one call's trace after 4,055 calls; its 11 games in flight were replayed in part 2 as the same attempt (DESIGN.md amendment of 2026-10-06; fixed at 76dfab0 for the later arms). Harness check, paired with E1's A2 on the same 460 games: A2 (Luna through OpenRouter) wins 8.7 points more on T2-T6 [4.4, 13.0], mostly on T5 and T6 (luna-codex-vs-a2.compare.md); under the plan's rule the harness changes play.
Exit code 0. Duration 9509 s.
- gpt-6-luna-codex-high.jsonl · experiments/E03-codex-frontier/results/runs/gpt-6-luna-codex-high.jsonl (run file; held by the Room owner; not public; ask the reporter) · restricted · sha256 17fbd1a92220… · 9216646 bytes · M41
- gpt-6-luna-codex-high.part2.jsonl · experiments/E03-codex-frontier/results/runs/gpt-6-luna-codex-high.part2.jsonl (run file; held by the Room owner; not public; ask the reporter) · restricted · sha256 088c3540fd64… · 9304489 bytes · M42
- gpt-6-luna-codex-high.traces.jsonl.gz · experiments/E03-codex-frontier/results/runs/traces/gpt-6-luna-codex-high.traces.jsonl.gz (call traces; held by the Room owner; not public; ask the reporter) · restricted · sha256 b0578c6f57cf… · 31811099 bytes · M43
- gpt-6-luna-codex-high.part2.traces.jsonl.gz · experiments/E03-codex-frontier/results/runs/traces/gpt-6-luna-codex-high.part2.traces.jsonl.gz (call traces; held by the Room owner; not public; ask the reporter) · restricted · sha256 010986519034… · 41045208 bytes · M44
- gpt-6-luna-codex-high.score.json · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/gpt-6-luna-codex-high.score.json · public · sha256 08fd09c39150… · 31958 bytes · M45
- gpt-6-luna-codex-high.diagnose.md · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/gpt-6-luna-codex-high.diagnose.md · public · sha256 0d70350ef5bf… · 1625 bytes · M46
- gpt-6-luna-codex-high.files.json · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/gpt-6-luna-codex-high.files.json · public · sha256 17356405ee4d… · 303 bytes · M47
- luna-codex-vs-a2.compare.json · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/luna-codex-vs-a2.compare.json · public · sha256 44c15636bab5… · 31019 bytes · M48
- luna-codex-vs-a2.compare.md · https://github.com/stw2/zendo-lab/blob/d351665aa/experiments/E03-codex-frontier/results/scores/luna-codex-vs-a2.compare.md · public · sha256 ec7c9d5ba3a3… · 3983 bytes · M49
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.