Execution attempt · A9 · planned
E2 arm qwen3.5-4b-mlx-show: Qwen3.5-4B shown the count of consistent rules after every observation, on the 102 dev games of the plan, per the accepted plan. Runs first, in one sequence on one Mac (show, force16, force16-show).
Pinned source and configuration
https://github.com/stw2/zendo-lab @ a028dc9c9e9aff82dd163bc5342a51610f6d62cd
Reference checked 2026-10-06 13:55 UTC. Later commits, branches or plan changes do not retarget this attempt.
- experiments/E02-small-model-interventions/DESIGN.md
- experiments/E02-small-model-interventions/scripts/01_run.sh
- experiments/E02-small-model-interventions/scripts/intervene.py
- experiments/E02-small-model-interventions/scripts/02_score.sh
- experiments/E02-small-model-interventions/scripts/03_measures.py
- pyproject.toml
- uv.lock
- Command
- bash experiments/E02-small-model-interventions/scripts/01_run.sh show
- Working directory
- .
- Configuration paths
- experiments/E02-small-model-interventions/scripts/intervene.py
- Parameters
- model=Qwen/Qwen3.5-4B@851bf6e8; backend=mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools, wrapper e02-intervention-v1; hardware=Apple M4 Max 128 GB; sampling=temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048; max_tokens=65536; force_experiments=None; show_consistent_rules=True
- Environment
- macOS 26.5 on Apple M4 Max 128 GB; Python 3.13 via uv; ZendoBench 1.0.0 MLX batch driver; mlx 0.32.2, mlx-lm 0.31.3
- Output directory
- experiments/E02-small-model-interventions/results/runs
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · dataset · Download
Delivered events
Registered
#1E2 arm qwen3.5-4b-mlx-show: Qwen3.5-4B shown the count of consistent rules after every observation, on the 102 dev games of the plan, per the accepted plan. Runs first, in one sequence on one Mac (show, force16, force16-show).
Started
#2Launched 2026-10-06T13:56:22Z at a028dc9 (clean tree): bash experiments/E02-small-model-interventions/scripts/01_run.sh show; 102 dev games (--per-family 2), MLX batch 36, sampler seed 0, player agent-tools; the run header records the wrapper's source sha256 (ded59be2…) and configuration (show_consistent_rules true).
Progress report
#3Interrupted and resumed, same attempt. The Mac lost power at about 2026-10-06T14:06Z, about 10 minutes into the run, before any game finished: qwen3.5-4b-mlx-show.jsonl holds its header and variant records and no episode, and the process left no exit record. Power returned at 19:39Z. At 19:57:33Z the run resumed per the plan with a second call of 01_run.sh show at a028dc9 (clean tree), into qwen3.5-4b-mlx-show.part2.jsonl with --exclude-finished over the first file (0 finished games), so all 102 games play in part 2. The configuration hash and the wrapper's source sha256 (ded59be2…) are unchanged. A10 and A11 follow it in the same sequence.
Error: power loss at about 2026-10-06T14:06Z; no exit record
Succeeded
#4Exit 0 at 2026-10-06T23:09:01Z; 102/102 games finished across the two files (part 1, cut by the power loss, holds none), and score --verify verified every game; 0 unscored, 0 malformed. T2-T6 win rate 4.9 [1.7, 21.5] over the 92 headline games (wins: T1 5/10, T2 2/14, T3 2/20, T4 0/10, T5 0/34, T6 0/14). Behaviour (diagnose, all 102 games): 2.30 experiments a game at 0.78 bits of expected information each; first submission with a median of 84.5 classes consistent (mean p_first 0.10); 29 of 201 submissions contradict the evidence. Run files restricted (sha256 given); summaries public at 33e9624. H3's test and the comparisons with A8 follow in E2's findings.
Exit code 0. Duration 11488 s.
- qwen3.5-4b-mlx-show.jsonl · experiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-show.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 7b1ff1e727a0… · 29288 bytes · M36
- qwen3.5-4b-mlx-show.part2.jsonl · experiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-show.part2.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 1a58a578fe3b… · 33433095 bytes · M37
- qwen3.5-4b-mlx-show.score.json · https://github.com/stw2/zendo-lab/blob/33e9624070bad7ef6c6ddba0c92812efbf7f2611/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-show.score.json · public · sha256 c734710309f4… · 24291 bytes · M38
- qwen3.5-4b-mlx-show.diagnose.md · https://github.com/stw2/zendo-lab/blob/33e9624070bad7ef6c6ddba0c92812efbf7f2611/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-show.diagnose.md · public · sha256 cb4eb7f66da7… · 1571 bytes · M39
- qwen3.5-4b-mlx-show.files.json · https://github.com/stw2/zendo-lab/blob/33e9624070bad7ef6c6ddba0c92812efbf7f2611/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-show.files.json · public · sha256 a3a0a92dea04… · 298 bytes · M40
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.