Experiment · E2
Is Qwen3.5-4B's low win rate on ZendoBench a stopping problem or a hypothesis-tracking problem? Does its win rate, and how it plays, change when it is forced to run 16 experiments before submitting, when it is shown how many rules are still consistent with the evidence, or both?
Qwen3.5-4B (bf16, revision 851bf6e8) on ZendoBench 1.0.0 dev, the first 2 items of each family (102 games; 92 T2-T6), with E1's model, sampling, hardware and seed. Three new arms, one run each: force16, show and force16-show, through a wrapper around ZendoBench's MLX batch engine (ZendoBench unpatched, player agent-tools). The no-intervention cell is E1's attempt A8 on the same games. Inference-time interventions only, no training; no train or sealed items.
Prerequisites and protocol
- Access needs
- Apple silicon with MLX (an M4 Max 128 GB here). E1's A8 run file is held by the owner (restricted), as these run files will be; scores, diagnostics and measures are public in the repository.
- Suggested protocol
- Pre-registered in experiments/E02-small-model-interventions/DESIGN.md (branch small-model-behaviour). Per arm: bash experiments/E02-small-model-interventions/scripts/01_run.sh ARM (force16, show, force16-show), then scripts/02_score.sh ARM (score --verify, diagnose), scripts/02_score.sh pairs (compare --verify), and scripts/03_measures.py (the paired comparisons with A8, the measures and the hypotheses' verdicts).
Accepted plan
Plan accepted by @stw2
Qwen3.5-4B on 102 ZendoBench 1.0.0 dev games (the first 2 items of each family) in three new arms, force16, show and force16-show, one run each, through a wrapper around ZendoBench's MLX batch engine. E1's A8 on the same games is the no-intervention cell. Scored with score --verify and diagnose; H1-H3 decided by scripts/03_measures.py.- requires zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · Download
- requires qwen3.5-4b-mlx.jsonl (E1 attempt A8) · baseline · M25 qwen3.5-4b-mlx.jsonl · Ask the reporter
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 3
3 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
planned attempt · stw2/zendo-lab @ a028dc9c9e9a · under an accepted plan
planned attempt · stw2/zendo-lab @ a028dc9c9e9a · under an accepted plan
planned attempt · stw2/zendo-lab @ a028dc9c9e9a · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find Learning to test hypotheses, thread Small models: a stopping problem or a hypothesis-tracking problem?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →