Finding

Sign in with GitHub
← Publications

Finding · P33 · Author-curated

On the 102 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit runs 16.25 experiments per game at 0.39 bits of expected information each, and first submits with a median of 2 rule classes still consistent with the evidence (an ideal player submitting then would win with mean probability 0.70; every game has a submission); 97 of its 189 submissions contradict evidence already shown, and it wins 19 of the 64.7 T2-T6 games such an ideal player would be expected to win (0.29). E2 attempt A10.

Published by @stw2 · 2026-10-08 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Behaviour profile

subject
Qwen3.5-4BThe model evaluated.
setting
E2 setting qwen3.5-4b-mlx-force16Backend, hardware, sampling, harness settings and intervention of the run.
intervention
Forced experimenting (16)What E2 changed against E1's setting.
evaluation items
ZendoBench 1.0.0 dev, 2 per familyThe games played.
experiments per game
16.25 experiments per gameValue of the concept experiments_per_game (P1).
expected information gain
0.388 bits per experimentValue of the concept expected_information_gain (P1).
classes alive at first submission
2 rule classesValue of the concept classes_alive_at_first_submission (P1; a median over games; for T6 counted within its analysis prior).
win probability at first submission
0.701 probabilityMean over games of diagnose's p_first: the probability that submitting the most probable rule class under the game's analysis prior wins at the first submission's moment (an ideal player's chance, not the submitted rule's; P1's concept of this name describes it otherwise).
contradicting submission share
0.513 share of submissionsValue of the concept contradicting_submission_share (defined in F1).
conversion to ideal wins
0.29 wins per expected ideal winValue of the concept conversion_to_ideal_wins (defined in F1).
preregistered
trueWhether E2's DESIGN.md named these measures before it was measured.

Experimental provenance

Method and evaluation protocol
E2: `zendo_bench diagnose` (ZendoBench 1.0.0) over the attempt's run files, from each game's engine record alone; conversion by scripts/03_measures.py.
Dataset
ZendoBench 1.0.0 dev manifest (Room material M1), its first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
exp_per_game 16.2549, bits_per_exp 0.4024, eig 0.3880, zero-information share 0.5078, repeats 86, alive_first 2.0, p_first 0.7006, submissions per game 1.8529, contradicting submissions 97/189, counterexample uptake 57/87, wrong submissions with exactly one class consistent 84, conversion 19/64.67.
Uncertainty and replication
Descriptive; no interval. Medians and means over the stated games.
Evidence references
qwen3.5-4b-mlx-force16.jsonl · restrictedexperiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-force16.jsonl (held by the Room owner; not public; ask the reporter)qwen3.5-4b-mlx-force16.diagnose.md · publichttps://github.com/stw2/zendo-lab/blob/9284883f59dfc88aac67d11519464f28c13570d0/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-force16.diagnose.mdmeasures.json · publichttps://github.com/stw2/zendo-lab/blob/5599dc15fe53b3613f4e85f6904b9cd43ba470a5/experiments/E02-small-model-interventions/results/measures.json
Limitations
One run per arm on 102 dev games (92 headline); one model; dev only, no sealed evaluation.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Behaviour profile

Predicate. Descriptive statistics of how the subject plays the evaluation items under the setting, from zendo_bench diagnose (ZendoBench 1.0.0): how much it experiments, the information its experiments gain, and how much uncertainty remains at its first submission. No interval.

Key behaviour_profile · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

ZendoBench 1.0.0 dev, 2 per family

Evaluation items. The first 2 items of each family of ZendoBench 1.0.0's dev manifest in manifest order (zendo_bench run --per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14, of which 92 form the T2-T6 headline. A subset of zendobench_1_0_0_dev (P1) in which every family of every tier is played. E2's games.

Key zendobench_1_0_0_dev_per_family_2 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Forced experimenting (16)

Intervention. Until the player has run 16 experiments, its observation's available_actions is ["experiment"] and its answer grammar is ZendoBench's own experiment-only variant; a sentence in the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions).

Key forced_experimenting_16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

E2 setting qwen3.5-4b-mlx-force16

E1's setting arm_qwen3_5_4b_mlx (P7: ZendoBench 1.0.0 MLX batch driver, Apple M4 Max 128 GB, bf16, 36 games in flight, sampler seed 0, Qwen3.5 thinking sampling, thinking budget 63,487 + 2,048-token answer allowance, answers constrained to valid actions) with forced_experimenting_16, applied by E2's wrapper e02-intervention-v1 (experiments/E02-small-model-interventions/scripts/intervene.py, sha256 ded59be23e3d236c3c63210803e7e3c1e0c4583663fd9c60c4345b84c275c4b6) after the game renders each request, ZendoBench unpatched; player agent-tools; `bash experiments/E02-small-model-interventions/scripts/01_run.sh force16`.

Key arm_qwen3_5_4b_mlx_force16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Qwen3.5-4B

Qwen3.5-4B open weights (bf16), Hugging Face Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, ZendoBench 1.0.0's pinned 4B checkpoint (shard SHA-256 checked on load).

Key qwen3_5_4b · version a1427e75-d571-45f4-9513-1276d692d84b

Exact references

related

b4fb1d38-4c66-4640-bfd9-d3f864fa4142

The headline estimate from the same attempt.

related

be062a4b-4e6a-4de7-be80-5e98a9ff5053

The same model without an intervention on the same games (A8).