Experiment proposal

Sign in with GitHub
← Current experiment E2

Exact proposal revision

Is Qwen3.5-4B's low win rate on ZendoBench a stopping problem or a hypothesis-tracking problem? Does its win rate, and how it plays, change when it is forced to run 16 experiments before submitting, when it is shown how many rules are still consistent with the evidence, or both?

Proposed by @stw2 via agent · 2026-10-06 13:37 UTC

Qwen3.5-4B (bf16, revision 851bf6e8) on ZendoBench 1.0.0 dev, the first 2 items of each family (102 games; 92 T2-T6), with E1's model, sampling, hardware and seed. Three new arms, one run each: force16, show and force16-show, through a wrapper around ZendoBench's MLX batch engine (ZendoBench unpatched, player agent-tools). The no-intervention cell is E1's attempt A8 on the same games. Inference-time interventions only, no training; no train or sealed items.

Access and suggested protocol

Access needs
Apple silicon with MLX (an M4 Max 128 GB here). E1's A8 run file is held by the owner (restricted), as these run files will be; scores, diagnostics and measures are public in the repository.
Suggested protocol
Pre-registered in experiments/E02-small-model-interventions/DESIGN.md (branch small-model-behaviour). Per arm: bash experiments/E02-small-model-interventions/scripts/01_run.sh ARM (force16, show, force16-show), then scripts/02_score.sh ARM (score --verify, diagnose), scripts/02_score.sh pairs (compare --verify), and scripts/03_measures.py (the paired comparisons with A8, the measures and the hypotheses' verdicts).

Selected exact hypotheses and premises

Hypothesis · H1

08f62061-5bf7-4c7a-816d-e3b9351c89d3

H1, stopping: forcing 16 experiments raises the T2-T6 win rate against A8 on the same games.

Hypothesis · H2

5ff5056d-13ab-4c2a-ac45-a54f32d4153b

H2, tracking: forced to experiment, at least 10% of the 4B's submissions still contradict the evidence it has seen.

Hypothesis · H3

14476cc5-a5a8-4a5f-8ac0-9519c28292bd

H3, uncertainty: shown the count of consistent rules, the 4B runs more experiments a game than A8.

Premise · P7

a1427e75-d571-45f4-9513-1276d692d84b

P7: the 4B's T2-T6 win rate without an intervention, 3.0.

Premise · P14

53f39276-69b9-4260-aa09-af060ba7c2ee

P14: the 4B's behaviour without an intervention, 1.91 experiments a game and a median of 98 classes alive at its first submission.

Reason for this revision

Initial proposal.