Attributed research checkpoint
E3 is complete: five OpenAI models through the Codex CLI with no tools, reasoning effort high, all 460 ZendoBench 1.0.0 dev games each, every game verified. Win rate on T2-T6: GPT-6 Astra 0.943 (P17), GPT-6.1 Sol 0.897 (P19), GPT-6 Sol 0.692 (P21), GPT-5.6 Terra 0.408 (P23), GPT-6 Luna 0.362 (P25); the seed-only baseline is 0.047 (P1). Behaviour (P18, P20, P22, P24, P26): Astra, 6.1 Sol and Sol run 17-19 experiments a game and first submit with a median of 1-2 rule classes alive; Terra runs 11.1 and submits with 3.5 alive; Luna runs 15.4 with 2 alive and makes no submission in 39 games. Harness check (P27): GPT-6 Luna wins 0.087 more through OpenRouter (E1 A2) than through Codex on the same games [0.044, 0.130], so E3's arms compare with each other, and with E1's arms only with that caveat.
Assignment, status and accepted plans remain authoritative on each experiment.
Reported state
- Reported progress
- E1 (seven systems over HTTP and MLX) and E3 (five systems through Codex) are complete on the same dev games. The strongest systems now win nearly every T2-T3 game; T4-T6 separate them. E2 (inference-time interventions on Qwen3.5-4B) is in progress in its own Thread.
- Open issues
- Why GPT-6 Luna plays worse through Codex is not identified: Codex's request (top_p 0.98, verbosity low, no output cap, ZendoBench's system prompt as Codex's instructions) or the snapshot OpenAI serves to Codex. The snapshots behind Codex's model slugs are neither pinned nor reported. One run per system, on dev. Astra and 6.1 Sol win at least 98% of T2-T3, so the dev headline separates them less than T4-T6 do.
- Suggested next action
- Evaluate more models in this Thread, as new experiments compared with E1 and E3 on the same dev games. The harness gap (P27) remains an open issue for any further Codex arm.
- Access needs
- A ChatGPT subscription with Codex access for any further Codex arm; for the campaign's training experiments, training compute and the owner's sealed evaluation.
Exact references
Exact record · P23
602a6b09-0683-4dc2-aebd-d042736f2922P23: Headline win rate of GPT-5.6 Terra (A15).
Exact record · P24
73127491-c89e-4685-bcad-48a9ff5e5cf4P24: Behaviour profile of GPT-5.6 Terra (A15).
Exact record · P27
f4ff537a-9cc4-4b5c-8550-ed3022a83291P27: Harness check: GPT-6 Luna through OpenRouter (E1 A2) against Codex (E3 A12).
Experiment · E3
Open experiment →Accepted plan
Exact plan →