Experiment · E1
Where do small and strong models stand on ZendoBench 1.0.0 dev, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?
ZendoBench 1.0.0, the dev manifest, all 460 games (T2-T6 headline, T1 diagnostics), one run per model, no training. Six models: Qwen3.5-2B (local MLX), Qwen3.5-27B-FP8 and Qwen3.8-27B-FP8 (vLLM 0.30, one H100 each), DeepSeek-V4-Flash-0731 (Together), gpt-oss-120b with reasoning high and GPT-6 Luna with reasoning high (OpenRouter). Every call capped at 65,536 output tokens; reasoning kept in the run files.
Prerequisites and protocol
- Access needs
- OpenRouter and Together API keys; two on-demand H100 sandboxes (Daytona) for the 27B models; Apple silicon for MLX. Run files are held by the owner (restricted); scores and diagnostics are public in the repository.
- Suggested protocol
- Per model: zendo_bench run --manifest dev (all items) with the model's backend and sampling, then score --verify and diagnose. One attempt per model; a stopped run resumes with --exclude-finished into a new file and the files are scored together.
Accepted plan
Plan accepted by @stw2
Seven models on the full ZendoBench 1.0.0 dev set (460 games), one run each, one output cap of 65,536 tokens a call, reasoning kept; scored with score --verify and described with diagnose. Revision: adds the qwen3.5-4b-mlx arm (DESIGN.md amendment of 2026-10-05); the six earlier arms and their attempts A1-A6 are unchanged.- requires zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · Download
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 8
1 preflight failed · 7 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
planned attempt · stw2/zendo-lab @ fac2c1b7381b · under an accepted plan
planned attempt · stw2/zendo-lab @ fac2c1b7381b · under an accepted plan
planned attempt · stw2/zendo-lab @ fac2c1b7381b · under an accepted plan
planned attempt · stw2/zendo-lab @ fac2c1b7381b · under an accepted plan
planned attempt · stw2/zendo-lab @ fac2c1b7381b · under an accepted plan
planned attempt · stw2/zendo-lab @ fac2c1b7381b · under an accepted plan
A7 E1 arm qwen3.5-4b-mlx: all 460 dev games, run per the accepted plan
Preflight failedplanned attempt · stw2/zendo-lab @ 1a5465257d33 · under an accepted plan
rerun attempt · stw2/zendo-lab @ 1a5465257d33 · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find Learning to test hypotheses, thread Campaign: can training teach small models to test hypotheses?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →