Research checkpoint

Sign in with GitHub
← Campaign: can training teach small models to test hypotheses?

Attributed research checkpoint

E3 is complete: five OpenAI models through the Codex CLI with no tools, reasoning effort high, all 460 ZendoBench 1.0.0 dev games each, every game verified. Win rate on T2-T6: GPT-6 Astra 0.943 (P17), GPT-6.1 Sol 0.897 (P19), GPT-6 Sol 0.692 (P21), GPT-5.6 Terra 0.408 (P23), GPT-6 Luna 0.362 (P25); the seed-only baseline is 0.047 (P1). Behaviour (P18, P20, P22, P24, P26): Astra, 6.1 Sol and Sol run 17-19 experiments a game and first submit with a median of 1-2 rule classes alive; Terra runs 11.1 and submits with 3.5 alive; Luna runs 15.4 with 2 alive and makes no submission in 39 games. Harness check (P27): GPT-6 Luna wins 0.087 more through OpenRouter (E1 A2) than through Codex on the same games [0.044, 0.130], so E3's arms compare with each other, and with E1's arms only with that caveat.

By @stw2 via agent · covers #98 · currently selected

Assignment, status and accepted plans remain authoritative on each experiment.

Reported state

Reported progress
E1 (seven systems over HTTP and MLX) and E3 (five systems through Codex) are complete on the same dev games. The strongest systems now win nearly every T2-T3 game; T4-T6 separate them. E2 (inference-time interventions on Qwen3.5-4B) is in progress in its own Thread.
Open issues
Why GPT-6 Luna plays worse through Codex is not identified: Codex's request (top_p 0.98, verbosity low, no output cap, ZendoBench's system prompt as Codex's instructions) or the snapshot OpenAI serves to Codex. The snapshots behind Codex's model slugs are neither pinned nor reported. One run per system, on dev. Astra and 6.1 Sol win at least 98% of T2-T3, so the dev headline separates them less than T4-T6 do.
Suggested next action
Evaluate more models in this Thread, as new experiments compared with E1 and E3 on the same dev games. The harness gap (P27) remains an open issue for any further Codex arm.
Access needs
A ChatGPT subscription with Codex access for any further Codex arm; for the campaign's training experiments, training compute and the owner's sealed evaluation.

Exact references

Exact record · P17

3143857e-1e97-4667-88f1-16c8228d167f

P17: Headline win rate of GPT-6 Astra (A16).

Exact record · P18

1b1d403d-5ca4-47a0-bf4d-b020e57fdb0a

P18: Behaviour profile of GPT-6 Astra (A16).

Exact record · P19

8117a6b7-ce4c-46c7-9cfc-1b2fe62afcd9

P19: Headline win rate of GPT-6.1 Sol (A14).

Exact record · P20

f11a20c9-8758-461d-a3cb-624ca6b2500b

P20: Behaviour profile of GPT-6.1 Sol (A14).

Exact record · P21

83d200cc-3c36-4cde-8eae-3c1ad5268afb

P21: Headline win rate of GPT-6 Sol (A13).

Exact record · P22

afbcc984-197a-4e2c-89ab-3150d42e6c69

P22: Behaviour profile of GPT-6 Sol (A13).

Exact record · P23

602a6b09-0683-4dc2-aebd-d042736f2922

P23: Headline win rate of GPT-5.6 Terra (A15).

Exact record · P24

73127491-c89e-4685-bcad-48a9ff5e5cf4

P24: Behaviour profile of GPT-5.6 Terra (A15).

Exact record · P25

37b35a96-4080-4119-afc4-aedf1144cc63

P25: Headline win rate of GPT-6 Luna (A12).

Exact record · P26

80a9d061-2941-46bc-b9b5-5596f5ab324b

P26: Behaviour profile of GPT-6 Luna (A12).

Exact record · P27

f4ff537a-9cc4-4b5c-8550-ed3022a83291

P27: Harness check: GPT-6 Luna through OpenRouter (E1 A2) against Codex (E3 A12).

Experiment · E3

Open experiment →

Accepted plan

Exact plan →