Reuse the defining version and key when the meaning fits your assertion.
zendo_bench compare 95% Student-t interval
The 95% Student-t interval that `zendo_bench compare` (ZendoBench 1.0.0) reports for a paired difference of equal-weight win rates; here over 215 rule classes with 66.5 degrees of freedom.
Key compare_student_t_95 · version f4ff537a-9cc4-4b5c-8550-ed3022a83291
Concept JSON
Paired difference in T2-T6 win rate
The difference in T2-T6 equal-weight win rate between two settings of the same model on the same games, paired by game, as `zendo_bench compare` (ZendoBench 1.0.0) computes it: setting_a minus setting_b.
Key paired_win_rate_difference · version f4ff537a-9cc4-4b5c-8550-ed3022a83291
Concept JSON
ZendoBench 1.0.0 dev
Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.
Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
Win rate, T2-T6 equal weight
Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.
Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
E1 setting gpt-6-luna-openrouter-high
HTTP chat completions to OpenRouter, default provider routing (the provider was not recorded per call), reasoning effort high, max_tokens 65,536, 64 calls in flight, reasoning text kept; player model, ZendoBench 1.0.0; run by `bash experiments/E01-dev-baseline/scripts/01_run.sh gpt-6-luna-openrouter-high` (github.com/stw2/zendo-lab).
Key arm_gpt_6_luna_openrouter_high · version 4afb1318-33ab-4d14-9f47-5bbeb4c6e769
Concept JSON · Defining publication
E3 setting gpt-6-luna-codex-high
The Codex CLI 0.160.0 (`codex exec`) with no tools in the request, one stateless call per decision (ZendoBench's system prompt as Codex's instructions, its user message as the prompt, the last agent message as the reply), reasoning effort high, reasoning summary detailed, Codex's defaults otherwise (temperature 1.0, top_p 0.98, verbosity low, no output cap), default service tier, 12 calls in flight, each call in a new empty folder under a macOS sandbox closing the repositories and run files; player model, ZendoBench 1.0.0; run by `bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-luna-codex-high` (github.com/stw2/zendo-lab).
Key arm_gpt_6_luna_codex_high · version 37b35a96-4080-4119-afc4-aedf1144cc63
Concept JSON · Defining publication
GPT-6 Luna
GPT-6 Luna (Codex model slug gpt-6-luna), a closed OpenAI model reached through the Codex CLI under a ChatGPT sign-in, accessed 2026-10-06; the version is whatever OpenAI served to Codex on those dates, neither pinned nor reported by the caller.
Key gpt_6_luna · version 37b35a96-4080-4119-afc4-aedf1144cc63
Concept JSON · Defining publication