Finding

Sign in with GitHub
← Publications

Finding · P38 · Author-curated

On the 102 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family), Qwen3.5-4B (MLX) shown after every observation how many rule classes are still consistent with the evidence runs 0.42 more experiments per game than without an intervention (E1's A8 on the same games): paired per-game difference, 95% CI 0.10-0.74; 2.30 against 1.88. E2 attempt A9 against E1 attempt A8.

Published by @stw2 · 2026-10-08 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired per-game mean difference

subject
Qwen3.5-4BThe model evaluated.
setting a
E2 setting qwen3.5-4b-mlx-showThe first setting (the difference is setting_a minus setting_b).
setting b
E1 setting qwen3.5-4b-mlxThe second setting.
intervention
Count of consistent rules shownThe intervention setting_a adds to setting_b.
evaluation items
ZendoBench 1.0.0 dev, 2 per familyThe games played.
metric
Experiments per gameThe quantity compared.
items
102 gamesGames scored in both settings (all tiers).
value
0.4216 experiments per gamePoint estimate of the difference.
interval low
0.1011 experiments per gameLower bound of the 95% interval.
interval high
0.742 experiments per gameUpper bound of the 95% interval.
interval method
Paired per-game Student-t 95% intervalHow the interval was computed.
preregistered
trueWhether E2's DESIGN.md named this comparison before it was measured.

Experimental provenance

Method and evaluation protocol
E2: scripts/03_measures.py pairs the two arms' games by task ID and takes the per-game difference in experiments (zendo_bench diagnose); both arms verified.
Dataset
ZendoBench 1.0.0 dev manifest (Room material M1), its first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
show - A8: 0.4216 [0.1011, 0.7420] experiments per game over 102 games; 50 games up, 28 down, 24 tied (two-sided sign test p = 0.017).
Uncertainty and replication
Student-t 95% interval on the per-game differences, 102 paired games, one run per arm. The lower bound is near 0 and E2 tests three hypotheses without a multiplicity correction; with a Bonferroni correction over three tests the interval is 0.03-0.81.
Evidence references
qwen3.5-4b-mlx-show.jsonl · restrictedexperiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-show.jsonl (held by the Room owner; not public; ask the reporter)qwen3.5-4b-mlx-show.part2.jsonl · restrictedexperiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-show.part2.jsonl (held by the Room owner; not public; ask the reporter)qwen3.5-4b-mlx-show.diagnose.md · publichttps://github.com/stw2/zendo-lab/blob/33e9624070bad7ef6c6ddba0c92812efbf7f2611/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-show.diagnose.mdqwen3.5-4b-mlx.jsonl · restrictedexperiments/E01-dev-baseline/results/runs/qwen3.5-4b-mlx.jsonl (held by the Room owner; not public; ask the reporter)measures.json · publichttps://github.com/stw2/zendo-lab/blob/5599dc15fe53b3613f4e85f6904b9cd43ba470a5/experiments/E02-small-model-interventions/results/measures.json
Limitations
One run per arm on 102 dev games (92 headline); one model; dev only, no sealed evaluation. The count also tells the model that its rule comes from a finite set of candidates. The no-intervention cell is E1's A8 restricted to these games, not a fresh control: a run of player model (a second attempt after an undecided verifier), from a 460-game run with 36 games in flight, on 2026-10-05; E2's arms are player agent-tools, 102 games, 36 in flight. The subset was fixed while A8's 0 of 92 headline wins on it was known; about 2.6 wins would be expected at A8's raw rate over all 430 headline games (12/430 = 0.028).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired per-game mean difference

Predicate. On the same games, the mean over games of a per-game measure under setting_a minus the same measure under setting_b, over the games scored under both, with a 95% interval.

Key paired_per_game_mean_difference · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

ZendoBench 1.0.0 dev, 2 per family

Evaluation items. The first 2 items of each family of ZendoBench 1.0.0's dev manifest in manifest order (zendo_bench run --per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14, of which 92 form the T2-T6 headline. A subset of zendobench_1_0_0_dev (P1) in which every family of every tier is played. E2's games.

Key zendobench_1_0_0_dev_per_family_2 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Paired per-game Student-t 95% interval

Interval method. Student-t 95% interval with n - 1 degrees of freedom on the per-game differences setting_a minus setting_b; E2's scripts/03_measures.py.

Key paired_t_95_per_game · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Count of consistent rules shown

Intervention. After every observation the player's observation carries consistent_rules {count, of}: of the classes of the game's analysis prior (1,765 for T1-T4, 619 for T5, 427 for T6), how many label every evidence scene shown (seeds, experiments, counterexamples) as shown, counted as zendo_bench diagnose counts classes alive, from the message alone and never from the hidden rule; a sentence in the system prompt explains the field and says the hidden rule is among those counted.

Key consistent_rules_shown · version 8cfab1fc-54aa-4522-8294-e0ec5f06fe4a

Experiments per game

Metric. Mean number of experiments (scenes the player built and had labelled) per scored game; zendo_bench diagnose field exp_per_game, over all scored dev games including T1.

Key experiments_per_game · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

E2 setting qwen3.5-4b-mlx-show

E1's setting arm_qwen3_5_4b_mlx (P7: ZendoBench 1.0.0 MLX batch driver, Apple M4 Max 128 GB, bf16, 36 games in flight, sampler seed 0, Qwen3.5 thinking sampling, thinking budget 63,487 + 2,048-token answer allowance, answers constrained to valid actions) with consistent_rules_shown, applied by E2's wrapper e02-intervention-v1 (experiments/E02-small-model-interventions/scripts/intervene.py, sha256 ded59be23e3d236c3c63210803e7e3c1e0c4583663fd9c60c4345b84c275c4b6) after the game renders each request, ZendoBench unpatched; player agent-tools; `bash experiments/E02-small-model-interventions/scripts/01_run.sh show`.

Key arm_qwen3_5_4b_mlx_show · version 8cfab1fc-54aa-4522-8294-e0ec5f06fe4a

E1 setting qwen3.5-4b-mlx

ZendoBench 1.0.0 MLX batch driver on Apple M4 Max 128 GB (mlx 0.32.2), bf16, 36 games in flight, sampler seed 0; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence penalty 1.5; thinking budget 63,487 tokens then a 2,048-token answer allowance (a call's cap 65,536), answers constrained to valid actions; reasoning kept; player model; `bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-4b-mlx`.

Key arm_qwen3_5_4b_mlx · version a1427e75-d571-45f4-9513-1276d692d84b

Qwen3.5-4B

Qwen3.5-4B open weights (bf16), Hugging Face Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, ZendoBench 1.0.0's pinned 4B checkpoint (shard SHA-256 checked on load).

Key qwen3_5_4b · version a1427e75-d571-45f4-9513-1276d692d84b

Exact references

supports

14476cc5-a5a8-4a5f-8ac0-9519c28292bd

H3's test: the paired per-game difference in experiments, show minus A8, has a 95% interval above 0.

derived from

d3ac9665-67e6-4f7f-af6c-086427122f93

The count arm's profile.

derived from

be062a4b-4e6a-4de7-be80-5e98a9ff5053

A8's profile on the same games.