Attributed research checkpoint
E1 (baseline on ZendoBench 1.0.0 dev) is complete. Win rate on T2-T6: GPT-6 Luna 0.449 (P2) and Qwen3.8-27B-FP8 0.445 (P3) lead; DeepSeek-V4-Flash 0.262 (P4), gpt-oss-120b 0.141 (P5), Qwen3.5-27B-FP8 0.104 (P6), Qwen3.5-4B 0.030 (P7), Qwen3.5-2B 0.000 (P8); the seed-only baseline is 0.047 (P1). Behaviour profiles P9-P15 (same order): the leaders run 14-17 experiments a game and first submit with a median 2 rule classes alive; Qwen3.5-2B runs 0.21 and submits with 188 alive. Across the seven systems, experiments per game and win rate rank identically (P16, descriptive). The campaign's premise holds on this benchmark: small models commit early with many rules alive.
Assignment, status and accepted plans remain authoritative on each experiment.
Reported state
- Reported progress
- Seven arms ran all 460 dev games with one output cap and reasoning kept; every game verified (A1-A6, A8; A7 failed its preflight and was rerun as A8). Sixteen findings follow one pattern: a headline estimate and a behaviour profile per system, with shared concepts defined in P1. Statistics, run-file digests and the record map are in github.com/stw2/zendo-lab, experiments/E01-dev-baseline (results/findings.json).
- Open issues
- Whether more experimenting causes more wins is untested: P16 compares seven systems that differ in size and training. Caveats: OpenRouter routed gpt-oss and Luna to providers that were not recorded; DeepSeek's headline is limited by the output cap; the MLX arms' answers are grammar-constrained, so malformed rates are not comparable across backends; one run per system, on dev.
- Suggested next action
- Propose E2: train Qwen3.5-2B or 4B on the train split with one or more of the campaign's signals (supervised traces from stronger players, RL on the win, a reward for the decision to stop), and compare with E1's attempts A1 and A8 as baselines on win rate, experiments per game and classes alive at the first submission.
- Access needs
- Training compute (GPU), the train split from zendo_bench train-sample, and the owner's sealed evaluation for final results.
Exact references
Exact record · P1
7f391bde-3e66-43b6-9a50-d299f7b7e5b2P1: seed-only baseline (0.047); defines the shared vocabulary for E1's records.
Exact record · P16
da8d956f-13f3-4dbc-b0e0-bb886394b401P16: rank agreement between experiments per game and win rate across the seven systems (descriptive).
Experiment · E1
Open experiment →Accepted plan
Exact plan →Accepted plan
Exact plan →