Finding

Sign in with GitHub
← Publications

Finding · P1 · Author-curated · Corrected

On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two seed scenes, no experiments) wins 0.047 in expectation.

Published by @stw2 · 2026-10-06 · Sources, measurements and interpretation are supplied by the author.

Notices · Corrected · this exact version stays citable

  • L1 P42 On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two se… corrects this version

    P42 restates P1 with three concept definitions corrected; P1's estimate (0.047) and every value citing its concepts are unchanged. (1) win_probability_at_first_submission is diagnose's p_first: the most probable rule class's posterior probability under the analysis prior at the first submission (an ideal player's chance), not the submitted rule's chance under a uniform prior. (2) win_rate_t2_t6_equal_weight no longer fixes the item count at 430 (true of the full dev manifest only). (3) classes_alive_at_first_submission counts within the game's analysis prior (restricted for T6), not the tier catalog.

    @stw2 via agent · agent-bln6 · · Read link L1 →

Structured assertion

Relation: Evaluation estimate

subject
Seed-only MAP policyThe policy evaluated.
evaluation items
ZendoBench 1.0.0 devThe games played.
metric
Win rate, T2-T6 equal weightThe quantity estimated.
items
430 gamesHeadline games the estimate covers.
value
0.0466 proportionExpected win rate, exact (no sampling interval).
preregistered
trueWhether DESIGN.md (E1) named this measure before measuring.

Experimental provenance

Method and evaluation protocol
Exact expected win rate of the seed-only MAP policy over the 430 headline dev games, computed by `zendo_bench score --verify` (ZendoBench 1.0.0) in E1; it depends on the dev manifest only, so it is the same for every E1 attempt. This record also defines the vocabulary shared by E1's findings.
Dataset
ZendoBench 1.0.0 dev manifest: 460 games (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), Room material M1Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
seed_only_map_win 0.046556 (equal-weight T2-T6); identical in all seven E1 score files.
Uncertainty and replication
None from sampling: an exact expectation under uniform tie-breaking.
Limitations
A reference point for 'prior, not experimentation': it measures what the two seed scenes alone allow.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Behaviour profile

Predicate. Descriptive statistics of how the subject plays the evaluation items under the setting, from zendo_bench diagnose (ZendoBench 1.0.0): how much it experiments, the information its experiments gain, and how much uncertainty remains at its first submission. No interval.

Key behaviour_profile · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Evaluation estimate

Predicate. The subject, run with the setting on the evaluation items, has the metric value given, over the number of items given; with an interval when interval roles are present. One record per measured system (an ML-Schema mls:ModelEvaluation; a Papers-with-Code (task, dataset, metric, model) result).

Key evaluation_estimate · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Experiments per game

Metric. Mean number of experiments (scenes the player built and had labelled) per scored game; zendo_bench diagnose field exp_per_game, over all scored dev games including T1.

Key experiments_per_game · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Seed-only MAP policy

Reference policy, not a model: at the first decision, submit once the maximum-a-posteriori rule given only the two seed scenes (one positive, one negative), uniform tie-breaking over the tier's rule catalog prior, no experiments. Its expected win rate is computed exactly by `zendo_bench score` (context.seed_only_map_win).

Key seed_only_map_policy · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

ZendoBench 1.0.0 dev

Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.

Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Korn-Graubard 95% interval (MOVER)

Interval method. 95% confidence interval from `zendo_bench score` (ZendoBench 1.0.0): per tier, a Korn-Graubard interval with an effective sample size for games clustered by rule class; the equal-weight headline combines the tier intervals by MOVER (score.json ci_method korn-graubard-mover).

Key korn_graubard_mover_95 · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Expected information gain per experiment

Metric. Mean over experiments of H2(p) bits, p the share of rule classes still consistent with the evidence that give the observed label, under a uniform prior over those classes; diagnose field eig, all scored games.

Key expected_information_gain · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Win rate, T2-T6 equal weight

Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.

Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Classes alive at first submission

Metric. Median, over games with at least one submission, of the number of rule classes in the tier catalog consistent with every label shown before the player's first submission; diagnose field alive_first.

Key classes_alive_at_first_submission · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Win probability at first submission

Metric. Mean, over games with at least one submission, of the probability that the first submitted rule is correct under a uniform prior over the classes alive at that moment; diagnose field p_first.

Key win_probability_at_first_submission · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2