Reuse the defining version and key when the meaning fits your assertion.
Paired per-game mean difference
Predicate. On the same games, the mean over games of a per-game measure under setting_a minus the same measure under setting_b, over the games scored under both, with a 95% interval.
Key paired_per_game_mean_difference · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142
Concept JSON · Defining publication
ZendoBench 1.0.0 dev, 2 per family
Evaluation items. The first 2 items of each family of ZendoBench 1.0.0's dev manifest in manifest order (zendo_bench run --per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14, of which 92 form the T2-T6 headline. A subset of zendobench_1_0_0_dev (P1) in which every family of every tier is played. E2's games.
Key zendobench_1_0_0_dev_per_family_2 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142
Concept JSON · Defining publication
Paired per-game Student-t 95% interval
Interval method. Student-t 95% interval with n - 1 degrees of freedom on the per-game differences setting_a minus setting_b; E2's scripts/03_measures.py.
Key paired_t_95_per_game · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142
Concept JSON · Defining publication
Count of consistent rules shown
Intervention. After every observation the player's observation carries consistent_rules {count, of}: of the classes of the game's analysis prior (1,765 for T1-T4, 619 for T5, 427 for T6), how many label every evidence scene shown (seeds, experiments, counterexamples) as shown, counted as zendo_bench diagnose counts classes alive, from the message alone and never from the hidden rule; a sentence in the system prompt explains the field and says the hidden rule is among those counted.
Key consistent_rules_shown · version 8cfab1fc-54aa-4522-8294-e0ec5f06fe4a
Concept JSON · Defining publication
Experiments per game
Metric. Mean number of experiments (scenes the player built and had labelled) per scored game; zendo_bench diagnose field exp_per_game, over all scored dev games including T1.
Key experiments_per_game · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
E2 setting qwen3.5-4b-mlx-show
E1's setting arm_qwen3_5_4b_mlx (P7: ZendoBench 1.0.0 MLX batch driver, Apple M4 Max 128 GB, bf16, 36 games in flight, sampler seed 0, Qwen3.5 thinking sampling, thinking budget 63,487 + 2,048-token answer allowance, answers constrained to valid actions) with consistent_rules_shown, applied by E2's wrapper e02-intervention-v1 (experiments/E02-small-model-interventions/scripts/intervene.py, sha256 ded59be23e3d236c3c63210803e7e3c1e0c4583663fd9c60c4345b84c275c4b6) after the game renders each request, ZendoBench unpatched; player agent-tools; `bash experiments/E02-small-model-interventions/scripts/01_run.sh show`.
Key arm_qwen3_5_4b_mlx_show · version 8cfab1fc-54aa-4522-8294-e0ec5f06fe4a
Concept JSON · Defining publication
E1 setting qwen3.5-4b-mlx
ZendoBench 1.0.0 MLX batch driver on Apple M4 Max 128 GB (mlx 0.32.2), bf16, 36 games in flight, sampler seed 0; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence penalty 1.5; thinking budget 63,487 tokens then a 2,048-token answer allowance (a call's cap 65,536), answers constrained to valid actions; reasoning kept; player model; `bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-4b-mlx`.
Key arm_qwen3_5_4b_mlx · version a1427e75-d571-45f4-9513-1276d692d84b
Concept JSON · Defining publication
Qwen3.5-4B
Qwen3.5-4B open weights (bf16), Hugging Face Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, ZendoBench 1.0.0's pinned 4B checkpoint (shard SHA-256 checked on load).
Key qwen3_5_4b · version a1427e75-d571-45f4-9513-1276d692d84b
Concept JSON · Defining publication