Execution attempt · A11 · planned
E2 arm qwen3.5-4b-mlx-force16-show: Qwen3.5-4B forced to run 16 experiments and shown the count of consistent rules, on the 102 dev games of the plan, per the accepted plan. Runs third, after force16, in one sequence on one Mac (show, force16, force16-show).
Pinned source and configuration
https://github.com/stw2/zendo-lab @ a028dc9c9e9aff82dd163bc5342a51610f6d62cd
Reference checked 2026-10-06 13:55 UTC. Later commits, branches or plan changes do not retarget this attempt.
- experiments/E02-small-model-interventions/DESIGN.md
- experiments/E02-small-model-interventions/scripts/01_run.sh
- experiments/E02-small-model-interventions/scripts/intervene.py
- experiments/E02-small-model-interventions/scripts/02_score.sh
- experiments/E02-small-model-interventions/scripts/03_measures.py
- pyproject.toml
- uv.lock
- Command
- bash experiments/E02-small-model-interventions/scripts/01_run.sh force16-show
- Working directory
- .
- Configuration paths
- experiments/E02-small-model-interventions/scripts/intervene.py
- Parameters
- model=Qwen/Qwen3.5-4B@851bf6e8; backend=mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, player agent-tools, wrapper e02-intervention-v1; hardware=Apple M4 Max 128 GB; sampling=temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048; max_tokens=65536; force_experiments=16; show_consistent_rules=True
- Environment
- macOS 26.5 on Apple M4 Max 128 GB; Python 3.13 via uv; ZendoBench 1.0.0 MLX batch driver; mlx 0.32.2, mlx-lm 0.31.3
- Output directory
- experiments/E02-small-model-interventions/results/runs
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · dataset · Download
Delivered events
Registered
#1E2 arm qwen3.5-4b-mlx-force16-show: Qwen3.5-4B forced to run 16 experiments and shown the count of consistent rules, on the 102 dev games of the plan, per the accepted plan. Runs third, after force16, in one sequence on one Mac (show, force16, force16-show).
Started
#2Launched 2026-10-07T16:10:12Z, right after A10 ended: bash experiments/E02-small-model-interventions/scripts/01_run.sh force16-show; 102 dev games (--per-family 2), MLX batch 36, sampler seed 0, player agent-tools; the run header records the wrapper's source sha256 (ded59be2…, as A9 and A10) and configuration (force_experiments 16, show_consistent_rules true). The working tree was clean at 33e9624, which differs from the registered a028dc9 only by A9's score summaries: the plan's source paths (DESIGN.md, scripts, pyproject.toml, uv.lock) are byte-identical.
Succeeded
#3Exit 0 at 2026-10-08T10:08:38Z; 102/102 games finished and score --verify verified every game; 0 unscored, 0 abandoned; malformed rate 0.1%. T2-T6 win rate 22.4 [13.6, 38.7] over the 92 headline games (wins: T1 8/10, T2 7/14, T3 10/20, T4 0/10, T5 4/34, T6 0/14). Behaviour (diagnose, all 102 games): 16.40 experiments a game at 0.39 bits of expected information each, 50% zero-information, 84 repeated scenes; first submission with a median of 2 classes consistent (mean p_first 0.67); 94 of 184 submissions contradict the evidence; 80 wrong submissions made with exactly one class consistent. compare --verify (verified): force16-show - force16 +3.3 points [-5.7, 12.2]; force16-show - show +17.5 [7.2, 27.8]. Run file restricted (sha256 given); summaries public at 370ce16. The comparisons with A8 and the hypotheses' tests follow in E2's findings.
Exit code 0. Duration 64706 s.
- qwen3.5-4b-mlx-force16-show.jsonl · experiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-force16-show.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 40e66401bf30… · 166197943 bytes · M79
- qwen3.5-4b-mlx-force16-show.score.json · https://github.com/stw2/zendo-lab/blob/370ce16dfa2fc09721ab2b8f606fa498f453d03d/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-force16-show.score.json · public · sha256 43c37f126dc3… · 22812 bytes · M80
- qwen3.5-4b-mlx-force16-show.diagnose.md · https://github.com/stw2/zendo-lab/blob/370ce16dfa2fc09721ab2b8f606fa498f453d03d/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-force16-show.diagnose.md · public · sha256 844ac37bdd19… · 1597 bytes · M81
- qwen3.5-4b-mlx-force16-show.files.json · https://github.com/stw2/zendo-lab/blob/370ce16dfa2fc09721ab2b8f606fa498f453d03d/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-force16-show.files.json · public · sha256 4588552a37c2… · 158 bytes · M82
- force16-show__force16.compare.json · https://github.com/stw2/zendo-lab/blob/370ce16dfa2fc09721ab2b8f606fa498f453d03d/experiments/E02-small-model-interventions/results/scores/force16-show__force16.compare.json · public · sha256 fa8f6a812910… · 22504 bytes · M83
- force16-show__show.compare.json · https://github.com/stw2/zendo-lab/blob/370ce16dfa2fc09721ab2b8f606fa498f453d03d/experiments/E02-small-model-interventions/results/scores/force16-show__show.compare.json · public · sha256 7412f773e04b… · 24218 bytes · M84
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.