Finding

Sign in with GitHub
← Publications

Finding · P21 · Author-curated

On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), GPT-6 Sol (reasoning high, Codex CLI, no tools) wins 0.692 (95% Korn-Graubard CI 0.646-0.735); seed-only baseline 0.047. E3 attempt A13.

Published by @stw2 · 2026-10-08 · Sources, measurements and interpretation are supplied by the author.

Advisory · cited versions changed

  • Cites a corrected version: P1 (related) · L1: P42 restates P1 with three concept definitions corrected; P1's estimate (0.047) and every value citing its concepts are unchanged. (1) win_probability_at_first_submission is diagnose's p_first: the most probable rule class's posterior probability under the analysis prior at the first submission (an ideal player's chance), not the submitted rule's chance under a uniform prior. (2) win_rate_t2_t6_equal_weight no longer fixes the item count at 430 (true of the full dev manifest only). (3) classes_alive_at_first_submission counts within the game's analysis prior (restricted for T6), not the tier catalog.

Structured assertion

Relation: Evaluation estimate

subject
GPT-6 SolThe model evaluated.
setting
E3 setting gpt-6-sol-codex-highHarness, sampling and isolation settings of the run.
evaluation items
ZendoBench 1.0.0 devThe games played.
metric
Win rate, T2-T6 equal weightThe quantity estimated.
items
430 gamesHeadline games finished and scored.
value
0.6920 proportionPoint estimate.
interval low
0.6462 proportionLower bound of the 95% interval.
interval high
0.7350 proportionUpper bound of the 95% interval.
interval method
Korn-Graubard 95% interval (MOVER)How the interval was computed.
comparator
Exact publication 7f391bde-3e66-43b6-9a50-d299f7b7e5b2The seed-only baseline on the same items.
preregistered
trueWhether DESIGN.md (E3) named this measure before measuring.

Experimental provenance

Method and evaluation protocol
E3 (pre-registered in experiments/E03-codex-frontier/DESIGN.md, github.com/stw2/zendo-lab): all 460 dev games played once with the setting above, scored by `zendo_bench score --verify`, which replays every game against its engine record.
Dataset
ZendoBench 1.0.0 dev manifest: 460 games (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), Room material M1Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
Headline 0.6920 [0.6462, 0.7350]; wins by tier T1 30/30; T2 58/60; T3 147/150; T4 17/60; T5 73/100; T6 30/60; malformed rate on headline decisions 0.0001.
Uncertainty and replication
95% Korn-Graubard intervals (MOVER for the equal-weight mean); one run, so sampling variation between runs is not measured.
Evidence references
gpt-6-sol-codex-high.jsonl · restrictedexperiments/E03-codex-frontier/results/runs/gpt-6-sol-codex-high.jsonl (run file; held by the Room owner; not public; ask the reporter)gpt-6-sol-codex-high.score.json · publichttps://github.com/stw2/zendo-lab/blob/664e563881be911512a4326d9d40ad6ef8630807/experiments/E03-codex-frontier/results/scores/gpt-6-sol-codex-high.score.json
Limitations
One run per system on the dev split, not sealed. Played through the Codex CLI with no tools at reasoning effort high: Codex sets no output cap and its own sampling defaults (temperature 1.0, top_p 0.98, verbosity low), and the model version is whatever OpenAI served to Codex on the dates stated, neither pinned nor reported. Through Codex, GPT-6 Luna wins 0.087 less than through OpenRouter on the same games (E3's harness check), so comparisons with E1's HTTP arms carry that difference. Dev rule catalogs are public; ZendoBench makes no contamination-free claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

GPT-6 Sol

GPT-6 Sol (Codex model slug gpt-6-sol), a closed OpenAI model reached through the Codex CLI under a ChatGPT sign-in, accessed 2026-10-06/07; the version is whatever OpenAI served to Codex on those dates, neither pinned nor reported by the caller.

Key gpt_6_sol · version 83d200cc-3c36-4cde-8eae-3c1ad5268afb

E3 setting gpt-6-sol-codex-high

The Codex CLI 0.160.0 (`codex exec`) with no tools in the request, one stateless call per decision (ZendoBench's system prompt as Codex's instructions, its user message as the prompt, the last agent message as the reply), reasoning effort high, reasoning summary detailed, Codex's defaults otherwise (temperature 1.0, top_p 0.98, verbosity low, no output cap), default service tier, 12 calls in flight, each call in a new empty folder under a macOS sandbox closing the repositories and run files; player model, ZendoBench 1.0.0; run by `bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-sol-codex-high` (github.com/stw2/zendo-lab).

Key arm_gpt_6_sol_codex_high · version 83d200cc-3c36-4cde-8eae-3c1ad5268afb

Evaluation estimate

Predicate. The subject, run with the setting on the evaluation items, has the metric value given, over the number of items given; with an interval when interval roles are present. One record per measured system (an ML-Schema mls:ModelEvaluation; a Papers-with-Code (task, dataset, metric, model) result).

Key evaluation_estimate · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

ZendoBench 1.0.0 dev

Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.

Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Korn-Graubard 95% interval (MOVER)

Interval method. 95% confidence interval from `zendo_bench score` (ZendoBench 1.0.0): per tier, a Korn-Graubard interval with an effective sample size for games clustered by rule class; the equal-weight headline combines the tier intervals by MOVER (score.json ci_method korn-graubard-mover).

Key korn_graubard_mover_95 · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Win rate, T2-T6 equal weight

Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.

Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Exact references

related

7f391bde-3e66-43b6-9a50-d299f7b7e5b2

The seed-only baseline on the same 430 games.