Finding

Sign in with GitHub
← Publications

Finding · P27 · Author-curated

On ZendoBench 1.0.0 dev (T2-T6, 430 games, paired, equal tier weights), GPT-6 Luna at reasoning effort high wins 0.087 more (95% CI 0.044-0.130) through OpenRouter (E1 attempt A2) than through the Codex CLI with no tools (E3 attempt A12); by tier, in points: T2 +0.0, T3 +8.0, T4 +3.3, T5 +19.0, T6 +13.3. Every Codex call returned a complete answer; which difference between the harnesses causes the gap is not identified.

Published by @stw2 · 2026-10-08 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference in T2-T6 win rate

subject
GPT-6 LunaThe model compared with itself.
setting a
E1 setting gpt-6-luna-openrouter-highThe first setting (the difference is setting_a minus setting_b).
setting b
E3 setting gpt-6-luna-codex-highThe second setting.
evaluation items
ZendoBench 1.0.0 devThe games played by both.
metric
Win rate, T2-T6 equal weightThe quantity compared.
items
430 gamesHeadline games finished and scored in both settings.
value
0.0873 proportionPoint estimate of the difference.
interval low
0.0445 proportionLower bound of the 95% interval.
interval high
0.1302 proportionUpper bound of the 95% interval.
interval method
zendo_bench compare 95% Student-t intervalHow the interval was computed.
preregistered
trueWhether DESIGN.md (E3) named this comparison before measuring.

Experimental provenance

Method and evaluation protocol
E3's pre-registered harness check (experiments/E03-codex-frontier/DESIGN.md): `zendo_bench compare --verify` of E1 attempt A2's run files (their sha256 checked against E1's committed files.json) with E3 attempt A12's, on the same 460 dev games (scripts/05_compare_a2.sh).
Dataset
ZendoBench 1.0.0 dev manifest: 460 games (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), Room material M1Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
Difference A2 minus A12 on T2-T6: 0.0873 [0.0445, 0.1302]; by tier, in points: T2 +0.0, T3 +8.0, T4 +3.3, T5 +19.0, T6 +13.3; T1 (diagnostic) -3.3. Every one of A12's 8,040 calls returned a complete answer and none was retried.
Uncertainty and replication
The 95% Student-t interval of the paired difference that zendo_bench compare reports; one run per setting.
Evidence references
luna-codex-vs-a2.compare.json · publichttps://github.com/stw2/zendo-lab/blob/d351665949c6109bbb6c78e7fe8084e95206193e/experiments/E03-codex-frontier/results/scores/luna-codex-vs-a2.compare.jsonluna-codex-vs-a2.compare.md · publichttps://github.com/stw2/zendo-lab/blob/d351665949c6109bbb6c78e7fe8084e95206193e/experiments/E03-codex-frontier/results/scores/luna-codex-vs-a2.compare.mdgpt-6-luna-codex-high.jsonl · restrictedexperiments/E03-codex-frontier/results/runs/gpt-6-luna-codex-high.jsonl (run file; held by the Room owner; not public; ask the reporter)gpt-6-luna-codex-high.part2.jsonl · restrictedexperiments/E03-codex-frontier/results/runs/gpt-6-luna-codex-high.part2.jsonl (run file; held by the Room owner; not public; ask the reporter)
Limitations
One run per setting, on dev. The two harnesses differ in request shape (chat completions through OpenRouter with default routing, the provider not recorded; the Responses API through Codex), in sampling defaults (Codex: top_p 0.98, verbosity low, no output cap; E1: max_tokens 65,536), in how the system prompt is passed, and possibly in the snapshot served; this comparison does not separate them.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

zendo_bench compare 95% Student-t interval

The 95% Student-t interval that `zendo_bench compare` (ZendoBench 1.0.0) reports for a paired difference of equal-weight win rates; here over 215 rule classes with 66.5 degrees of freedom.

Key compare_student_t_95 · version f4ff537a-9cc4-4b5c-8550-ed3022a83291

Paired difference in T2-T6 win rate

The difference in T2-T6 equal-weight win rate between two settings of the same model on the same games, paired by game, as `zendo_bench compare` (ZendoBench 1.0.0) computes it: setting_a minus setting_b.

Key paired_win_rate_difference · version f4ff537a-9cc4-4b5c-8550-ed3022a83291

ZendoBench 1.0.0 dev

Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.

Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Win rate, T2-T6 equal weight

Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.

Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

E1 setting gpt-6-luna-openrouter-high

HTTP chat completions to OpenRouter, default provider routing (the provider was not recorded per call), reasoning effort high, max_tokens 65,536, 64 calls in flight, reasoning text kept; player model, ZendoBench 1.0.0; run by `bash experiments/E01-dev-baseline/scripts/01_run.sh gpt-6-luna-openrouter-high` (github.com/stw2/zendo-lab).

Key arm_gpt_6_luna_openrouter_high · version 4afb1318-33ab-4d14-9f47-5bbeb4c6e769

E3 setting gpt-6-luna-codex-high

The Codex CLI 0.160.0 (`codex exec`) with no tools in the request, one stateless call per decision (ZendoBench's system prompt as Codex's instructions, its user message as the prompt, the last agent message as the reply), reasoning effort high, reasoning summary detailed, Codex's defaults otherwise (temperature 1.0, top_p 0.98, verbosity low, no output cap), default service tier, 12 calls in flight, each call in a new empty folder under a macOS sandbox closing the repositories and run files; player model, ZendoBench 1.0.0; run by `bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-luna-codex-high` (github.com/stw2/zendo-lab).

Key arm_gpt_6_luna_codex_high · version 37b35a96-4080-4119-afc4-aedf1144cc63

GPT-6 Luna

GPT-6 Luna (Codex model slug gpt-6-luna), a closed OpenAI model reached through the Codex CLI under a ChatGPT sign-in, accessed 2026-10-06; the version is whatever OpenAI served to Codex on those dates, neither pinned nor reported by the caller.

Key gpt_6_luna · version 37b35a96-4080-4119-afc4-aedf1144cc63

Exact references

related

4afb1318-33ab-4d14-9f47-5bbeb4c6e769

GPT-6 Luna through OpenRouter (E1 A2): the comparison's first setting.

related

37b35a96-4080-4119-afc4-aedf1144cc63

GPT-6 Luna through Codex (E3 A12): the comparison's second setting.