Finding

Sign in with GitHub
← Publications

Finding · P16 · Author-curated

Among the seven systems evaluated in E1 on ZendoBench 1.0.0 dev, ranking by experiments per game matches ranking by T2-T6 win rate exactly (Spearman rho 1.00, 7 systems). Descriptive of these systems only; not a causal or population claim.

Published by @stw2 · 2026-10-06 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Rank agreement among systems

x metric
Experiments per gameFirst measure ranked.
y metric
Win rate, T2-T6 equal weightSecond measure ranked.
evaluation items
ZendoBench 1.0.0 devThe games behind both measures.
systems
7 systemsSystems ranked (E1's seven evaluation pairs, cited in references).
statistic
Spearman's rhoThe rank statistic.
value
1 rhoSpearman's rho, ties at average ranks.
preregistered
falseWhether DESIGN.md (E1) named this analysis before measuring.

Experimental provenance

Method and evaluation protocol
Spearman's rho over E1's seven systems between each behaviour measure and the T2-T6 headline, by experiments/E01-dev-baseline/scripts/04_rank_agreement.py from results/metrics.json (github.com/stw2/zendo-lab).
Dataset
ZendoBench 1.0.0 dev manifest: 460 games (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), Room material M1Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
Spearman rho with the headline: experiments per game 1.00, classes alive at first submission -0.991, win probability at first submission 1.00, expected information gain -0.25.
Uncertainty and replication
No significance test: the seven systems are not a random sample, and both measures come from the same games.
Limitations
Seven heterogeneous systems that differ in size, provider, training and settings; capability is a common cause of both measures, so the agreement says nothing about whether more experimenting causes more wins. Not pre-registered.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Spearman's rho

Pearson correlation of the ranks, ties at average ranks; no significance test reported.

Key spearman_rho · version da8d956f-13f3-4dbc-b0e0-bb886394b401

Rank agreement among systems

Predicate. Over the listed systems, ordering them by the x_metric and by the y_metric gives the rank statistic shown. Describes these systems only: no claim about a population of systems, about a causal effect, or about variation within one system.

Key rank_agreement_among_systems · version da8d956f-13f3-4dbc-b0e0-bb886394b401

ZendoBench 1.0.0 dev

Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.

Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Experiments per game

Metric. Mean number of experiments (scenes the player built and had labelled) per scored game; zendo_bench diagnose field exp_per_game, over all scored dev games including T1.

Key experiments_per_game · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Win rate, T2-T6 equal weight

Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.

Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Exact references

derived from

4afb1318-33ab-4d14-9f47-5bbeb4c6e769

y_metric of A2.

derived from

9f72911f-8a38-4906-90c6-2bb0d3637236

y_metric of A5.

derived from

f4d4ccd3-23dc-4b95-b06f-788b853ee24f

y_metric of A4.

derived from

4abd87e2-3745-4e85-8857-70705d976a05

y_metric of A3.

derived from

c7096cbf-a84e-4c7d-b719-6fea1e79ca98

y_metric of A6.

derived from

a1427e75-d571-45f4-9513-1276d692d84b

y_metric of A8.

derived from

e3552371-0135-40d4-ac20-412b614ce6f2

y_metric of A1.

derived from

14fbeacc-5178-44b1-aa46-0593126c8f6a

x_metric of A2.

derived from

da204c2b-aed5-4f7c-ae35-56745141f111

x_metric of A5.

derived from

40b19c65-c17d-44ab-97f8-b73331ad4443

x_metric of A4.

derived from

a68c0e7e-0f62-4632-9cf5-8c098b7f4e1d

x_metric of A3.

derived from

f164a7c1-d061-4a5f-83ea-e6980601af35

x_metric of A6.

derived from

53f39276-69b9-4260-aa09-af060ba7c2ee

x_metric of A8.

derived from

cd29dc79-e65d-4baa-a036-42920fbc748b

x_metric of A1.