Exact proposal revision
Where do small and strong models stand on ZendoBench 1.0.0 dev, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?
ZendoBench 1.0.0, the dev manifest, all 460 games (T2-T6 headline, T1 diagnostics), one run per model, no training. Six models: Qwen3.5-2B (local MLX), Qwen3.5-27B-FP8 and Qwen3.8-27B-FP8 (vLLM 0.30, one H100 each), DeepSeek-V4-Flash-0731 (Together), gpt-oss-120b with reasoning high and GPT-6 Luna with reasoning high (OpenRouter). Every call capped at 65,536 output tokens; reasoning kept in the run files.
Access and suggested protocol
- Access needs
- OpenRouter and Together API keys; two on-demand H100 sandboxes (Daytona) for the 27B models; Apple silicon for MLX. Run files are held by the owner (restricted); scores and diagnostics are public in the repository.
- Suggested protocol
- Per model: zendo_bench run --manifest dev (all items) with the model's backend and sampling, then score --verify and diagnose. One attempt per model; a stopped run resumes with --exclude-finished into a new file and the files are scored together.
Selected exact hypotheses and premises
No hypotheses or premises selected.
Reason for this revision
Initial proposal.