Exact proposal revision
Is Qwen3.5-4B's low win rate on ZendoBench a stopping problem or a hypothesis-tracking problem? Does its win rate, and how it plays, change when it is forced to run 16 experiments before submitting, when it is shown how many rules are still consistent with the evidence, or both?
Qwen3.5-4B (bf16, revision 851bf6e8) on ZendoBench 1.0.0 dev, the first 2 items of each family (102 games; 92 T2-T6), with E1's model, sampling, hardware and seed. Three new arms, one run each: force16, show and force16-show, through a wrapper around ZendoBench's MLX batch engine (ZendoBench unpatched, player agent-tools). The no-intervention cell is E1's attempt A8 on the same games. Inference-time interventions only, no training; no train or sealed items.
Access and suggested protocol
- Access needs
- Apple silicon with MLX (an M4 Max 128 GB here). E1's A8 run file is held by the owner (restricted), as these run files will be; scores, diagnostics and measures are public in the repository.
- Suggested protocol
- Pre-registered in experiments/E02-small-model-interventions/DESIGN.md (branch small-model-behaviour). Per arm: bash experiments/E02-small-model-interventions/scripts/01_run.sh ARM (force16, show, force16-show), then scripts/02_score.sh ARM (score --verify, diagnose), scripts/02_score.sh pairs (compare --verify), and scripts/03_measures.py (the paired comparisons with A8, the measures and the hypotheses' verdicts).
Selected exact hypotheses and premises
Hypothesis · H1
08f62061-5bf7-4c7a-816d-e3b9351c89d3H1, stopping: forcing 16 experiments raises the T2-T6 win rate against A8 on the same games.
Hypothesis · H2
5ff5056d-13ab-4c2a-ac45-a54f32d4153bH2, tracking: forced to experiment, at least 10% of the 4B's submissions still contradict the evidence it has seen.
Hypothesis · H3
14476cc5-a5a8-4a5f-8ac0-9519c28292bdH3, uncertainty: shown the count of consistent rules, the 4B runs more experiments a game than A8.
Premise · P7
a1427e75-d571-45f4-9513-1276d692d84bP7: the 4B's T2-T6 win rate without an intervention, 3.0.
Premise · P14
53f39276-69b9-4260-aa09-af060ba7c2eeP14: the 4B's behaviour without an intervention, 1.91 experiments a game and a median of 98 classes alive at its first submission.
Reason for this revision
Initial proposal.