Finding

Sign in with GitHub
← Publications

Finding · P36 · Author-curated

On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit has a T2-T6 win rate 0.191 higher than without an intervention (E1's A8 on the same games): paired difference, 95% CI 0.095-0.287. The forced arm first submitted with a median of 2 consistent rule classes, so the precondition for reading H1 holds. E2 attempt A10 against E1 attempt A8.

Published by @stw2 · 2026-10-08 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference in T2-T6 win rate

subject
Qwen3.5-4BThe model evaluated.
setting a
E2 setting qwen3.5-4b-mlx-force16The first setting (the difference is setting_a minus setting_b).
setting b
E1 setting qwen3.5-4b-mlxThe second setting.
intervention
Forced experimenting (16)The intervention setting_a adds to setting_b.
evaluation items
ZendoBench 1.0.0 dev, 2 per familyThe games played.
metric
Win rate, T2-T6 equal weightThe quantity compared.
items
92 gamesHeadline games finished and scored in both settings.
value
0.191 proportionPoint estimate of the difference.
interval low
0.0949 proportionLower bound of the 95% interval.
interval high
0.2871 proportionUpper bound of the 95% interval.
interval method
zendo_bench compare 95% Student-t intervalHow the interval was computed.
preregistered
trueWhether E2's DESIGN.md named this comparison before it was measured.

Experimental provenance

Method and evaluation protocol
E2: E1's A8 is a run of player model, which `zendo_bench compare` refuses to pair with E2's agent-tools runs, so scripts/03_measures.py applies compare's pairing by task ID and its breakdown (the same statistics) to the two arms' scored games; E2's arm verified by score --verify, A8 in E1.
Dataset
ZendoBench 1.0.0 dev manifest (Room material M1), its first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
force16 - A8: 0.1910 [0.0949, 0.2871] over 92 paired headline games. force16 0.1910, A8 0.0000.
Uncertainty and replication
`zendo_bench compare`'s Student-t 95% interval on rule-class means (families as strata, equal tier weights); 92 paired headline games, one run per arm. The interval method is P27's compare_student_t_95; the class and degree-of-freedom counts in P27's definition are P27's own comparison's, not this one's.
Evidence references
qwen3.5-4b-mlx-force16.jsonl · restrictedexperiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-force16.jsonl (held by the Room owner; not public; ask the reporter)qwen3.5-4b-mlx-force16.score.json · publichttps://github.com/stw2/zendo-lab/blob/9284883f59dfc88aac67d11519464f28c13570d0/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-force16.score.jsonqwen3.5-4b-mlx.jsonl · restrictedexperiments/E01-dev-baseline/results/runs/qwen3.5-4b-mlx.jsonl (held by the Room owner; not public; ask the reporter)measures.json · publichttps://github.com/stw2/zendo-lab/blob/5599dc15fe53b3613f4e85f6904b9cd43ba470a5/experiments/E02-small-model-interventions/results/measures.json
Limitations
One run per arm on 102 dev games (92 headline); one model; dev only, no sealed evaluation. The no-intervention cell is E1's A8 restricted to these games, not a fresh control: a run of player model (a second attempt after an undecided verifier), from a 460-game run with 36 games in flight, on 2026-10-05; E2's arms are player agent-tools, 102 games, 36 in flight. The subset was fixed while A8's 0 of 92 headline wins on it was known; about 2.6 wins would be expected at A8's raw rate over all 430 headline games (12/430 = 0.028).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference in T2-T6 win rate

The difference in T2-T6 equal-weight win rate between two settings of the same model on the same games, paired by game, as `zendo_bench compare` (ZendoBench 1.0.0) computes it: setting_a minus setting_b.

Key paired_win_rate_difference · version f4ff537a-9cc4-4b5c-8550-ed3022a83291

ZendoBench 1.0.0 dev, 2 per family

Evaluation items. The first 2 items of each family of ZendoBench 1.0.0's dev manifest in manifest order (zendo_bench run --per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14, of which 92 form the T2-T6 headline. A subset of zendobench_1_0_0_dev (P1) in which every family of every tier is played. E2's games.

Key zendobench_1_0_0_dev_per_family_2 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

zendo_bench compare 95% Student-t interval

The 95% Student-t interval that `zendo_bench compare` (ZendoBench 1.0.0) reports for a paired difference of equal-weight win rates; here over 215 rule classes with 66.5 degrees of freedom.

Key compare_student_t_95 · version f4ff537a-9cc4-4b5c-8550-ed3022a83291

Forced experimenting (16)

Intervention. Until the player has run 16 experiments, its observation's available_actions is ["experiment"] and its answer grammar is ZendoBench's own experiment-only variant; a sentence in the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions).

Key forced_experimenting_16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Win rate, T2-T6 equal weight

Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.

Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

E2 setting qwen3.5-4b-mlx-force16

E1's setting arm_qwen3_5_4b_mlx (P7: ZendoBench 1.0.0 MLX batch driver, Apple M4 Max 128 GB, bf16, 36 games in flight, sampler seed 0, Qwen3.5 thinking sampling, thinking budget 63,487 + 2,048-token answer allowance, answers constrained to valid actions) with forced_experimenting_16, applied by E2's wrapper e02-intervention-v1 (experiments/E02-small-model-interventions/scripts/intervene.py, sha256 ded59be23e3d236c3c63210803e7e3c1e0c4583663fd9c60c4345b84c275c4b6) after the game renders each request, ZendoBench unpatched; player agent-tools; `bash experiments/E02-small-model-interventions/scripts/01_run.sh force16`.

Key arm_qwen3_5_4b_mlx_force16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

E1 setting qwen3.5-4b-mlx

ZendoBench 1.0.0 MLX batch driver on Apple M4 Max 128 GB (mlx 0.32.2), bf16, 36 games in flight, sampler seed 0; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence penalty 1.5; thinking budget 63,487 tokens then a 2,048-token answer allowance (a call's cap 65,536), answers constrained to valid actions; reasoning kept; player model; `bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-4b-mlx`.

Key arm_qwen3_5_4b_mlx · version a1427e75-d571-45f4-9513-1276d692d84b

Qwen3.5-4B

Qwen3.5-4B open weights (bf16), Hugging Face Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, ZendoBench 1.0.0's pinned 4B checkpoint (shard SHA-256 checked on load).

Key qwen3_5_4b · version a1427e75-d571-45f4-9513-1276d692d84b

Exact references

supports

08f62061-5bf7-4c7a-816d-e3b9351c89d3

H1's test: the paired difference force16 minus A8 has a 95% interval above 0, and the precondition holds (median 2 classes alive at the first submission, at most 10).

derived from

4a45cc3f-4aa0-4f61-8c88-164a0a483b54

The forced arm's profile, for the precondition.

derived from

9c6080af-4e48-4962-82be-e9d303ce9238

A8's estimate on the same games.

derived from

b4fb1d38-4c66-4640-bfd9-d3f864fa4142

The forced arm's estimate.

related

f4ff537a-9cc4-4b5c-8550-ed3022a83291

Defines the paired win-rate difference frame and compare's interval, used here.