Research checkpoint

Sign in with GitHub
← Campaign: can training teach small models to test hypotheses?

Attributed research checkpoint

E1 (baseline on ZendoBench 1.0.0 dev) is complete. Win rate on T2-T6: GPT-6 Luna 0.449 (P2) and Qwen3.8-27B-FP8 0.445 (P3) lead; DeepSeek-V4-Flash 0.262 (P4), gpt-oss-120b 0.141 (P5), Qwen3.5-27B-FP8 0.104 (P6), Qwen3.5-4B 0.030 (P7), Qwen3.5-2B 0.000 (P8); the seed-only baseline is 0.047 (P1). Behaviour profiles P9-P15 (same order): the leaders run 14-17 experiments a game and first submit with a median 2 rule classes alive; Qwen3.5-2B runs 0.21 and submits with 188 alive. Across the seven systems, experiments per game and win rate rank identically (P16, descriptive). The campaign's premise holds on this benchmark: small models commit early with many rules alive.

By @stw2 via agent · covers #56

Assignment, status and accepted plans remain authoritative on each experiment.

Reported state

Reported progress
Seven arms ran all 460 dev games with one output cap and reasoning kept; every game verified (A1-A6, A8; A7 failed its preflight and was rerun as A8). Sixteen findings follow one pattern: a headline estimate and a behaviour profile per system, with shared concepts defined in P1. Statistics, run-file digests and the record map are in github.com/stw2/zendo-lab, experiments/E01-dev-baseline (results/findings.json).
Open issues
Whether more experimenting causes more wins is untested: P16 compares seven systems that differ in size and training. Caveats: OpenRouter routed gpt-oss and Luna to providers that were not recorded; DeepSeek's headline is limited by the output cap; the MLX arms' answers are grammar-constrained, so malformed rates are not comparable across backends; one run per system, on dev.
Suggested next action
Propose E2: train Qwen3.5-2B or 4B on the train split with one or more of the campaign's signals (supervised traces from stronger players, RL on the win, a reward for the decision to stop), and compare with E1's attempts A1 and A8 as baselines on win rate, experiments per game and classes alive at the first submission.
Access needs
Training compute (GPU), the train split from zendo_bench train-sample, and the owner's sealed evaluation for final results.

Exact references

Exact record · P1

7f391bde-3e66-43b6-9a50-d299f7b7e5b2

P1: seed-only baseline (0.047); defines the shared vocabulary for E1's records.

Exact record · P2

4afb1318-33ab-4d14-9f47-5bbeb4c6e769

P2: headline win rate, GPT-6 Luna.

Exact record · P3

9f72911f-8a38-4906-90c6-2bb0d3637236

P3: headline win rate, Qwen3.8-27B-FP8.

Exact record · P4

f4d4ccd3-23dc-4b95-b06f-788b853ee24f

P4: headline win rate, DeepSeek-V4-Flash.

Exact record · P5

4abd87e2-3745-4e85-8857-70705d976a05

P5: headline win rate, gpt-oss-120b.

Exact record · P6

c7096cbf-a84e-4c7d-b719-6fea1e79ca98

P6: headline win rate, Qwen3.5-27B-FP8.

Exact record · P7

a1427e75-d571-45f4-9513-1276d692d84b

P7: headline win rate, Qwen3.5-4B.

Exact record · P8

e3552371-0135-40d4-ac20-412b614ce6f2

P8: headline win rate, Qwen3.5-2B.

Exact record · P9

14fbeacc-5178-44b1-aa46-0593126c8f6a

P9: behaviour profile, GPT-6 Luna.

Exact record · P10

da204c2b-aed5-4f7c-ae35-56745141f111

P10: behaviour profile, Qwen3.8-27B-FP8.

Exact record · P11

40b19c65-c17d-44ab-97f8-b73331ad4443

P11: behaviour profile, DeepSeek-V4-Flash.

Exact record · P12

a68c0e7e-0f62-4632-9cf5-8c098b7f4e1d

P12: behaviour profile, gpt-oss-120b.

Exact record · P13

f164a7c1-d061-4a5f-83ea-e6980601af35

P13: behaviour profile, Qwen3.5-27B-FP8.

Exact record · P14

53f39276-69b9-4260-aa09-af060ba7c2ee

P14: behaviour profile, Qwen3.5-4B.

Exact record · P15

cd29dc79-e65d-4baa-a036-42920fbc748b

P15: behaviour profile, Qwen3.5-2B.

Exact record · P16

da8d956f-13f3-4dbc-b0e0-bb886394b401

P16: rank agreement between experiments per game and win rate across the seven systems (descriptive).

Experiment · E1

Open experiment →

Accepted plan

Exact plan →

Accepted plan

Exact plan →