- Method and evaluation protocol
- E3: `zendo_bench diagnose` (ZendoBench 1.0.0) over the attempt's run files, from each game's engine record alone.
- Dataset
- ZendoBench 1.0.0 dev manifest: 460 games (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), Room material M1Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
- Reported results
- exp_per_game 17.1370, bits_per_exp 0.4071, eig 0.3910, zero-information share 0.5443, alive_first 1, p_first 0.7005, submissions per game 1.4370, contradicting submissions 2/661, counterexample uptake 201/201, games with a submission 460/460.
- Uncertainty and replication
- Descriptive; no interval. Medians and means over the stated games.
- Limitations
- One run per system on the dev split, not sealed. Played through the Codex CLI with no tools at reasoning effort high: Codex sets no output cap and its own sampling defaults (temperature 1.0, top_p 0.98, verbosity low), and the model version is whatever OpenAI served to Codex on the dates stated, neither pinned nor reported. Through Codex, GPT-6 Luna wins 0.087 less than through OpenRouter on the same games (E3's harness check), so comparisons with E1's HTTP arms carry that difference. Dev rule catalogs are public; ZendoBench makes no contamination-free claim. Covers all 460 games including T1, while the headline covers T2-T6.