Accepted plan

Sign in with GitHub
← Experiment E2

Immutable accepted plan · prospective

Qwen3.5-4B on 102 ZendoBench 1.0.0 dev games (the first 2 items of each family) in three new arms, force16, show and force16-show, one run each, through a wrapper around ZendoBench's MLX batch engine. E1's A8 on the same games is the no-intervention cell. Scored with score --verify and diagnose; H1-H3 decided by scripts/03_measures.py.

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

First plan for E2: the design pre-registered in DESIGN.md at commit a028dc9, with A8 as the no-intervention baseline.

Public source

Plan

Prediction
H1 (stopping): force16 raises the T2-T6 win rate against A8, paired 95% interval above 0. H2 (tracking): at least 10% of force16's submissions contradict the evidence. H3 (uncertainty): show raises experiments per game against A8, paired 95% interval above 0. They are predictions under test; the design commits to no outcome.
Protocol
Pre-registered in experiments/E02-small-model-interventions/DESIGN.md at this commit. Before the runs, scripts/00_check.py plays one game per tier in each arm with a scripted stand-in (every file must pass score --verify), and SMOKE=1 plays one T1 game of force16-show on the 4B; neither is a measurement. Per arm: `bash experiments/E02-small-model-interventions/scripts/01_run.sh ARM` from the repository root (ARM: force16, show, force16-show) writes results/runs/qwen3.5-4b-mlx-ARM.jsonl through scripts/intervene.py: ZendoBench's own MLX setup with E1's 4B arguments, --per-family 2, player agent-tools. A stopped run resumes with a second call (--exclude-finished into .partN.jsonl) and stays the same attempt. Then `bash .../scripts/02_score.sh ARM` (score --verify, diagnose, the run files' sha256), `bash .../scripts/02_score.sh pairs` (compare --verify force16-show against force16 and show) and `uv run python .../scripts/03_measures.py` (the paired comparisons with A8 on the same games, by task ID with compare's pairing and statistics; the measures; H1-H3's verdicts). One attempt per arm; run files are restricted outputs reported with their sha256 and size, summaries public at their commit. Order: show, force16, force16-show.
Dataset
ZendoBench 1.0.0 dev manifest (zendo_bench/data/bench-v1-dev.json, file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2, Room material M1), from github.com/stw2/zendo-bench tag v1.0.0 (commit 46c192e0ea10a5140a33c1280edb97b0127cc68c). Baseline: E1 attempt A8's run file (M25, restricted).
Split
dev, the first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14; T2-T6 (92) form the headline, T1 is diagnostics. No train or sealed items.
Access needs
Apple silicon with MLX. The run files and A8's run file are held by the owner (restricted); scores, diagnostics and measures are public in the repository.
Configurations
qwen3.5-4b-mlx-force16: Qwen/Qwen3.5-4B@851bf6e8, mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; wrapper e02-intervention-v1 with force_experiments=16, show_consistent_rules=False; qwen3.5-4b-mlx-show: Qwen/Qwen3.5-4B@851bf6e8, mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; wrapper e02-intervention-v1 with force_experiments=None, show_consistent_rules=True; qwen3.5-4b-mlx-force16-show: Qwen/Qwen3.5-4B@851bf6e8, mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; wrapper e02-intervention-v1 with force_experiments=16, show_consistent_rules=True
Metric
Headline: equal-weight T2-T6 win rate with 95% CI (score --verify). Paired: the T2-T6 win difference and the per-game experiments difference against A8 (force16, show) and between new arms (force16-show against force16 and show). Diagnose: experiments per game, bits per experiment, EIG, zero-information share, repeats, classes alive and p at the first submission, submissions per game, certain-but-wrong, contradicting share, counterexample uptake; conversion (T2-T6 wins over the summed p_first).
Seeds
Episode seeds from the dev manifest. MLX sampler seed 0 (each call's seed a hash of the session seed, task, attempt and decision, as in A8).
Interpretation rule
Precondition for H1 and H2: force16's median classes alive at the first submission is at most 10; otherwise neither is read and the result is about choosing experiments. H1 holds if the paired difference force16 - A8 has a 95% interval above 0; H2 holds if force16's contradicting share is at least 0.10; H3 holds if the paired experiments difference show - A8 has a 95% interval above 0. H1 holds and H2 fails: a stopping problem; H2 holds and H1 fails: a tracking problem; both: both; neither: read conversion and the experiment measures. An interval including 0 is reported as no difference measured.
Resources
Local MLX on an Apple M4 Max 128 GB: roughly 3 h for show and 12-15 h for each forced arm.
Prior work
E1 (A8) on all 460 dev games; on these 102 games A8 won 0 of 92 T2-T6 games, ran 1.88 experiments a game, submitted first with a median of 90 classes alive, and 21.7% of its submissions contradicted the evidence. Before the plan: 00_check.py passed for every arm, and one smoke game (T1, force16-show) passed score --verify with the gate held for 16 experiments; neither is a measurement or reported.

Selected exact hypotheses and premises

Hypothesis · H1

08f62061-5bf7-4c7a-816d-e3b9351c89d3

H1, stopping: decided by the paired win difference force16 - A8.

Hypothesis · H2

5ff5056d-13ab-4c2a-ac45-a54f32d4153b

H2, tracking: decided by force16's share of contradicting submissions.

Hypothesis · H3

14476cc5-a5a8-4a5f-8ac0-9519c28292bd

H3, uncertainty: decided by the paired experiments-per-game difference show - A8.

Premise · P7

a1427e75-d571-45f4-9513-1276d692d84b

P7: the 4B's T2-T6 win rate without an intervention.

Premise · P14

53f39276-69b9-4260-aa09-af060ba7c2ee

P14: the 4B's behaviour without an intervention (A8).