- Method and evaluation protocol
- E2: E1's A8 run file restricted to E2's 102 games by scripts/03_measures.py, scored with `zendo_bench score`'s breakdown and described with `zendo_bench diagnose`; A8's games were verified in E1. Conversion by scripts/03_measures.py.
- Dataset
- ZendoBench 1.0.0 dev manifest (Room material M1), its first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
- Reported results
- exp_per_game 1.8824, bits_per_exp 0.8336, eig 0.7558, zero-information share 0.0677, repeats 1, alive_first 90.0, p_first 0.0848, submissions per game 1.9902, contradicting submissions 44/203, counterexample uptake 91/101, wrong submissions with exactly one class consistent 5, conversion 0/7.60.
- Uncertainty and replication
- Descriptive; no interval. Medians and means over the stated games.
- Limitations
- One run per arm on 102 dev games (92 headline); one model; dev only, no sealed evaluation. The no-intervention cell is E1's A8 restricted to these games, not a fresh control: a run of player model (a second attempt after an undecided verifier), from a 460-game run with 36 games in flight, on 2026-10-05; E2's arms are player agent-tools, 102 games, 36 in flight. The subset was fixed while A8's 0 of 92 headline wins on it was known; about 2.6 wins would be expected at A8's raw rate over all 430 headline games (12/430 = 0.028).