Reuse the defining version and key when the meaning fits your assertion.
GPT-6 Astra
GPT-6 Astra (Codex model slug gpt-6-astra), a closed OpenAI model reached through the Codex CLI under a ChatGPT sign-in, accessed 2026-10-07; the version is whatever OpenAI served to Codex on those dates, neither pinned nor reported by the caller.
Key gpt_6_astra · version 3143857e-1e97-4667-88f1-16c8228d167f
Concept JSON
E3 setting gpt-6-astra-codex-high
The Codex CLI 0.160.0 (`codex exec`) with no tools in the request, one stateless call per decision (ZendoBench's system prompt as Codex's instructions, its user message as the prompt, the last agent message as the reply), reasoning effort high, reasoning summary detailed, Codex's defaults otherwise (temperature 1.0, top_p 0.98, verbosity low, no output cap), default service tier, 12 calls in flight, each call in a new empty folder under a macOS sandbox closing the repositories and run files; player model, ZendoBench 1.0.0; run by `bash experiments/E03-codex-frontier/scripts/01_run.sh gpt-6-astra-codex-high` (github.com/stw2/zendo-lab).
Key arm_gpt_6_astra_codex_high · version 3143857e-1e97-4667-88f1-16c8228d167f
Concept JSON
Evaluation estimate
Predicate. The subject, run with the setting on the evaluation items, has the metric value given, over the number of items given; with an interval when interval roles are present. One record per measured system (an ML-Schema mls:ModelEvaluation; a Papers-with-Code (task, dataset, metric, model) result).
Key evaluation_estimate · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
ZendoBench 1.0.0 dev
Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.
Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
Korn-Graubard 95% interval (MOVER)
Interval method. 95% confidence interval from `zendo_bench score` (ZendoBench 1.0.0): per tier, a Korn-Graubard interval with an effective sample size for games clustered by rule class; the equal-weight headline combines the tier intervals by MOVER (score.json ci_method korn-graubard-mover).
Key korn_graubard_mover_95 · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
Win rate, T2-T6 equal weight
Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.
Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication