Immutable accepted plan · prospective
Seven models on the full ZendoBench 1.0.0 dev set (460 games), one run each, one output cap of 65,536 tokens a call, reasoning kept; scored with score --verify and described with diagnose. Revision: adds the qwen3.5-4b-mlx arm (DESIGN.md amendment of 2026-10-05); the six earlier arms and their attempts A1-A6 are unchanged.
The owner asked for a seventh arm, Qwen3.5-4B on MLX, after A1-A6 had started; it is added before any 4B game, with the amendment committed at 1a54652. Earlier arms unchanged.
Public source
https://github.com/stw2/zendo-lab @ 1a5465257d3325741ecf2a7abdfada884cfa5e49
Reference checked 2026-10-05 10:07 UTC. No code was executed or scientific result verified.
- experiments/E01-dev-baseline/DESIGN.md
- experiments/E01-dev-baseline/scripts/01_run.sh
- experiments/E01-dev-baseline/scripts/02_score.sh
- experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json
- experiments/E01-dev-baseline/scripts/daytona/ctl.py
- experiments/E01-dev-baseline/scripts/daytona/guard.py
- experiments/E01-dev-baseline/scripts/daytona/setup.sh
- experiments/E01-dev-baseline/scripts/daytona/serve.sh
- pyproject.toml
- uv.lock
Plan
- Prediction
- Descriptive baseline, no hypothesis. Expected from the campaign's premise and a pre-release pilot: Qwen3.5-2B near zero, submitting with many classes alive; the 27B and frontier arms well above it.
- Protocol
- Pre-registered in experiments/E01-dev-baseline/DESIGN.md at this commit. Per arm: `bash experiments/E01-dev-baseline/scripts/01_run.sh ARM` from the repository root writes experiments/E01-dev-baseline/results/runs/ARM.jsonl (all 460 dev items, player model, the arm's backend and sampling as listed in DESIGN.md and the script). The two Qwen 27B arms are served by vLLM 0.30.0 on one on-demand H100 each, set up by experiments/E01-dev-baseline/scripts/daytona/ctl.py (model files checked against the Hub's hashes, then one short chat completion) and watched by guard.py (stops at $200 or 34 h). Before a run, SMOKE=1 plays one game per tier as a connection check; it is not a measurement. A stopped run resumes with a second call (--exclude-finished into ARM.partN.jsonl) and stays the same attempt. Then `bash experiments/E01-dev-baseline/scripts/02_score.sh ARM`: score --verify and diagnose over the arm's files, the run files' sha256. One attempt per arm; run files are reported as restricted outputs with their sha256 and size, the score and diagnose summaries as public outputs at their commit. Revision of 2026-10-05: a seventh arm, qwen3.5-4b-mlx, added before any 4B game was played (DESIGN.md amendments, commit 1a54652).
- Dataset
- ZendoBench 1.0.0 dev manifest (zendo_bench/data/bench-v1-dev.json, file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2, manifest hash 6b98d1a5...), from github.com/stw2/zendo-bench tag v1.0.0 (commit 46c192e0ea10a5140a33c1280edb97b0127cc68c).
- Split
- dev, all 460 items: T1 30, T2 60, T3 150, T4 60, T5 100, T6 60. T2-T6 form the headline; T1 is diagnostics. No train or sealed items.
- Access needs
- OpenRouter and Together API keys; two on-demand Daytona H100 sandboxes; Apple silicon with MLX. Run files are held by the owner (restricted); scores and diagnostics are public in the repository.
- Configurations
- qwen3.5-4b-mlx: Qwen/Qwen3.5-4B@851bf6e8, mlx batch driver, 36 games in flight, sampler seed 0, --checkpoint 4B, temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; qwen3.5-2b-mlx: Qwen/Qwen3.5-2B@15852e8c, mlx batch driver, 36 games in flight, sampler seed 0, temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; thinking budget 63487 + answer allowance 2048, max_tokens 65536; qwen3.5-27b-fp8-vllm-h100: Qwen/Qwen3.5-27B-FP8@97f5941b, vLLM 0.30.0, --batch auto, temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5, max_tokens 65536; qwen3.8-27b-fp8-vllm-h100: Qwen/Qwen3.8-27B-FP8@017b9c7a, vLLM 0.30.0, --batch auto, temperature 1.0, top_p 0.95, top_k 20, presence_penalty 0, reasoning_effort xhigh, max_tokens 65536; deepseek-v4-flash-together: deepseek-ai/DeepSeek-V4-Flash-0731, Together serverless, 64 calls in flight, provider defaults, max_tokens 65536; gpt-oss-120b-openrouter-high: openai/gpt-oss-120b, OpenRouter, default routing, 64 calls in flight, reasoning effort high, max_tokens 65536; gpt-6-luna-openrouter-high: openai/gpt-6-luna, OpenRouter, 64 calls in flight, reasoning effort high, max_tokens 65536
- Metric
- Headline: equal-weight mean win rate over T2-T6 with 95% CI (score --verify). Secondary: wins per tier, T1, unscored and unfinished counts, malformed rate, seed-only baseline and w0 bands; diagnose: experiments per game, bits per experiment, expected information gain, zero-information share, classes alive and p at the first submission, rule parsimony, contradictions, counterexample uptake; tokens, calls cut at the cap, wall time, list cost.
- Seeds
- Episode seeds from the dev manifest. MLX sampler seed 0. APIs and vLLM unseeded.
- Interpretation rule
- Descriptive. Arms are compared only where 95% CIs separate. The campaign premise fails on this benchmark if Qwen3.5-2B submits first with few classes alive (median alive_first at most 2, or p_first at least 0.5) or its headline lies inside a 27B arm's CI.
- Resources
- Daytona: two on-demand H100s, about 21-30 h each, about $125-180 each at list price, capped at $200 / 34 h each. APIs: about $40-50 in total. MLX: local, hours to days. The 4B arm: local MLX, hours.
- Prior work
- A pre-release pilot (36 dev games per model, 2026-10-04, ZendoBench commit d51b81d) sized the budget; its numbers are not results and are not reused.
Selected exact hypotheses and premises
No hypotheses or premises selected.