Finding

Sign in with GitHub
← Publications

Finding · P28 · Author-curated

On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit has a T2-T6 win rate of 0.191 (95% Korn-Graubard CI 0.114-0.348); the seed-only baseline on the same games is 0.033. E2 attempt A10.

Published by @stw2 · 2026-10-08 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Evaluation estimate

subject
Qwen3.5-4BThe model evaluated.
setting
E2 setting qwen3.5-4b-mlx-force16Backend, hardware, sampling, harness settings and intervention of the run.
intervention
Forced experimenting (16)What E2 changed against E1's setting.
evaluation items
ZendoBench 1.0.0 dev, 2 per familyThe games played.
metric
Win rate, T2-T6 equal weightThe quantity estimated.
items
92 gamesHeadline games finished and scored.
value
0.191 proportionPoint estimate.
interval low
0.1143 proportionLower bound of the 95% interval.
interval high
0.3478 proportionUpper bound of the 95% interval.
interval method
Korn-Graubard 95% interval (MOVER)How the interval was computed.
preregistered
trueWhether E2's DESIGN.md named this measure before it was measured.

Experimental provenance

Method and evaluation protocol
E2: `zendo_bench score --verify` (ZendoBench 1.0.0) over the attempt's run files; every game replayed from its recorded replies.
Dataset
ZendoBench 1.0.0 dev manifest (Room material M1), its first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
T2-T6 0.1910 [0.1143, 0.3478]; wins T1 7/10, T2 5/14, T3 7/20, T4 0/10, T5 6/34, T6 1/14; 0 unscored; every game verified. Seed-only baseline on the same games 0.033. Pooled w0 bands: w0=0 0.199 (45 classes); w0>=0.5 0.000 (3 classes; too few for an interval). Malformed rate (headline) 0.0006. 1848 calls; 0 reached the thinking budget (none was cut at the 65,536-token cap); 1 ended inside the thought without an answer.
Uncertainty and replication
95% Korn-Graubard intervals per tier with rule classes as clusters, combined over T2-T6 by MOVER; 92 headline games, one run.
Evidence references
qwen3.5-4b-mlx-force16.jsonl · restrictedexperiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-force16.jsonl (held by the Room owner; not public; ask the reporter)qwen3.5-4b-mlx-force16.score.json · publichttps://github.com/stw2/zendo-lab/blob/9284883f59dfc88aac67d11519464f28c13570d0/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-force16.score.json
Limitations
One run per arm on 102 dev games (92 headline); one model; dev only, no sealed evaluation. The wrapper changes the prompt after the game renders it, so the run is labelled player agent-tools.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired per-game Student-t 95% interval

Interval method. Student-t 95% interval with n - 1 degrees of freedom on the per-game differences setting_a minus setting_b; E2's scripts/03_measures.py.

Key paired_t_95_per_game · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Forced experimenting (16)

Intervention. Until the player has run 16 experiments, its observation's available_actions is ["experiment"] and its answer grammar is ZendoBench's own experiment-only variant; a sentence in the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions).

Key forced_experimenting_16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Conversion to ideal wins

Metric. Among the T2-T6 games with a submission, the games won divided by the sum of zendo_bench diagnose's p_first: the probability that submitting the most probable rule class under the game's analysis prior wins at the first submission's moment. Wins relative to those an ideal player submitting at the same moments would expect. E2's scripts/03_measures.py.

Key conversion_to_ideal_wins · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

E2 setting qwen3.5-4b-mlx-force16

E1's setting arm_qwen3_5_4b_mlx (P7: ZendoBench 1.0.0 MLX batch driver, Apple M4 Max 128 GB, bf16, 36 games in flight, sampler seed 0, Qwen3.5 thinking sampling, thinking budget 63,487 + 2,048-token answer allowance, answers constrained to valid actions) with forced_experimenting_16, applied by E2's wrapper e02-intervention-v1 (experiments/E02-small-model-interventions/scripts/intervene.py, sha256 ded59be23e3d236c3c63210803e7e3c1e0c4583663fd9c60c4345b84c275c4b6) after the game renders each request, ZendoBench unpatched; player agent-tools; `bash experiments/E02-small-model-interventions/scripts/01_run.sh force16`.

Key arm_qwen3_5_4b_mlx_force16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Contradicting submission share

Metric. The share of a player's submissions whose rule mislabels at least one evidence scene shown before it (seeds, experiments, counterexamples); zendo_bench diagnose field contradicting (n of m submissions), over all scored games including T1.

Key contradicting_submission_share · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Paired per-game mean difference

Predicate. On the same games, the mean over games of a per-game measure under setting_a minus the same measure under setting_b, over the games scored under both, with a 95% interval.

Key paired_per_game_mean_difference · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

ZendoBench 1.0.0 dev, 2 per family

Evaluation items. The first 2 items of each family of ZendoBench 1.0.0's dev manifest in manifest order (zendo_bench run --per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14, of which 92 form the T2-T6 headline. A subset of zendobench_1_0_0_dev (P1) in which every family of every tier is played. E2's games.

Key zendobench_1_0_0_dev_per_family_2 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Evaluation estimate

Predicate. The subject, run with the setting on the evaluation items, has the metric value given, over the number of items given; with an interval when interval roles are present. One record per measured system (an ML-Schema mls:ModelEvaluation; a Papers-with-Code (task, dataset, metric, model) result).

Key evaluation_estimate · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Korn-Graubard 95% interval (MOVER)

Interval method. 95% confidence interval from `zendo_bench score` (ZendoBench 1.0.0): per tier, a Korn-Graubard interval with an effective sample size for games clustered by rule class; the equal-weight headline combines the tier intervals by MOVER (score.json ci_method korn-graubard-mover).

Key korn_graubard_mover_95 · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Win rate, T2-T6 equal weight

Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.

Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

Qwen3.5-4B

Qwen3.5-4B open weights (bf16), Hugging Face Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, ZendoBench 1.0.0's pinned 4B checkpoint (shard SHA-256 checked on load).

Key qwen3_5_4b · version a1427e75-d571-45f4-9513-1276d692d84b

Exact references

related

a1427e75-d571-45f4-9513-1276d692d84b

E1's estimate for the same model and base setting, on all 430 headline games.