Immutable accepted plan · prospective
Qwen3.5-4B on 102 ZendoBench 1.0.0 dev games (the first 2 items of each family) in three new arms, force16, show and force16-show, one run each, through a wrapper around ZendoBench's MLX batch engine. E1's A8 on the same games is the no-intervention cell. Scored with score --verify and diagnose; H1-H3 decided by scripts/03_measures.py.
First plan for E2: the design pre-registered in DESIGN.md at commit a028dc9, with A8 as the no-intervention baseline.
Public source
https://github.com/stw2/zendo-lab @ a028dc9c9e9aff82dd163bc5342a51610f6d62cd
Reference checked 2026-10-06 13:55 UTC. No code was executed or scientific result verified.
- experiments/E02-small-model-interventions/DESIGN.md
- experiments/E02-small-model-interventions/scripts/00_check.py
- experiments/E02-small-model-interventions/scripts/01_run.sh
- experiments/E02-small-model-interventions/scripts/02_score.sh
- experiments/E02-small-model-interventions/scripts/03_measures.py
- experiments/E02-small-model-interventions/scripts/intervene.py
- pyproject.toml
- uv.lock
Plan
- Prediction
- H1 (stopping): force16 raises the T2-T6 win rate against A8, paired 95% interval above 0. H2 (tracking): at least 10% of force16's submissions contradict the evidence. H3 (uncertainty): show raises experiments per game against A8, paired 95% interval above 0. They are predictions under test; the design commits to no outcome.
- Protocol
- Pre-registered in experiments/E02-small-model-interventions/DESIGN.md at this commit. Before the runs, scripts/00_check.py plays one game per tier in each arm with a scripted stand-in (every file must pass score --verify), and SMOKE=1 plays one T1 game of force16-show on the 4B; neither is a measurement. Per arm: `bash experiments/E02-small-model-interventions/scripts/01_run.sh ARM` from the repository root (ARM: force16, show, force16-show) writes results/runs/qwen3.5-4b-mlx-ARM.jsonl through scripts/intervene.py: ZendoBench's own MLX setup with E1's 4B arguments, --per-family 2, player agent-tools. A stopped run resumes with a second call (--exclude-finished into .partN.jsonl) and stays the same attempt. Then `bash .../scripts/02_score.sh ARM` (score --verify, diagnose, the run files' sha256), `bash .../scripts/02_score.sh pairs` (compare --verify force16-show against force16 and show) and `uv run python .../scripts/03_measures.py` (the paired comparisons with A8 on the same games, by task ID with compare's pairing and statistics; the measures; H1-H3's verdicts). One attempt per arm; run files are restricted outputs reported with their sha256 and size, summaries public at their commit. Order: show, force16, force16-show.
- Dataset
- ZendoBench 1.0.0 dev manifest (zendo_bench/data/bench-v1-dev.json, file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2, Room material M1), from github.com/stw2/zendo-bench tag v1.0.0 (commit 46c192e0ea10a5140a33c1280edb97b0127cc68c). Baseline: E1 attempt A8's run file (M25, restricted).
- Split
- dev, the first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14; T2-T6 (92) form the headline, T1 is diagnostics. No train or sealed items.
- Access needs
- Apple silicon with MLX. The run files and A8's run file are held by the owner (restricted); scores, diagnostics and measures are public in the repository.
- Configurations
- qwen3.5-4b-mlx-force16: Qwen/Qwen3.5-4B@851bf6e8, mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; wrapper e02-intervention-v1 with force_experiments=16, show_consistent_rules=False; qwen3.5-4b-mlx-show: Qwen/Qwen3.5-4B@851bf6e8, mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; wrapper e02-intervention-v1 with force_experiments=None, show_consistent_rules=True; qwen3.5-4b-mlx-force16-show: Qwen/Qwen3.5-4B@851bf6e8, mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; wrapper e02-intervention-v1 with force_experiments=16, show_consistent_rules=True
- Metric
- Headline: equal-weight T2-T6 win rate with 95% CI (score --verify). Paired: the T2-T6 win difference and the per-game experiments difference against A8 (force16, show) and between new arms (force16-show against force16 and show). Diagnose: experiments per game, bits per experiment, EIG, zero-information share, repeats, classes alive and p at the first submission, submissions per game, certain-but-wrong, contradicting share, counterexample uptake; conversion (T2-T6 wins over the summed p_first).
- Seeds
- Episode seeds from the dev manifest. MLX sampler seed 0 (each call's seed a hash of the session seed, task, attempt and decision, as in A8).
- Interpretation rule
- Precondition for H1 and H2: force16's median classes alive at the first submission is at most 10; otherwise neither is read and the result is about choosing experiments. H1 holds if the paired difference force16 - A8 has a 95% interval above 0; H2 holds if force16's contradicting share is at least 0.10; H3 holds if the paired experiments difference show - A8 has a 95% interval above 0. H1 holds and H2 fails: a stopping problem; H2 holds and H1 fails: a tracking problem; both: both; neither: read conversion and the experiment measures. An interval including 0 is reported as no difference measured.
- Resources
- Local MLX on an Apple M4 Max 128 GB: roughly 3 h for show and 12-15 h for each forced arm.
- Prior work
- E1 (A8) on all 460 dev games; on these 102 games A8 won 0 of 92 T2-T6 games, ran 1.88 experiments a game, submitted first with a median of 90 classes alive, and 21.7% of its submissions contradicted the evidence. Before the plan: 00_check.py passed for every arm, and one smoke game (T1, force16-show) passed score --verify with the gate held for 16 experiments; neither is a measurement or reported.
Selected exact hypotheses and premises
Hypothesis · H1
08f62061-5bf7-4c7a-816d-e3b9351c89d3H1, stopping: decided by the paired win difference force16 - A8.
Hypothesis · H2
5ff5056d-13ab-4c2a-ac45-a54f32d4153bH2, tracking: decided by force16's share of contradicting submissions.
Hypothesis · H3
14476cc5-a5a8-4a5f-8ac0-9519c28292bdH3, uncertainty: decided by the paired experiments-per-game difference show - A8.
Premise · P7
a1427e75-d571-45f4-9513-1276d692d84bP7: the 4B's T2-T6 win rate without an intervention.
Premise · P14
53f39276-69b9-4260-aa09-af060ba7c2eeP14: the 4B's behaviour without an intervention (A8).