Research checkpoint

Sign in with GitHub
← Small models: a stopping problem or a hypothesis-tracking problem?

Attributed research checkpoint

E2 is complete: A9-A11 against E1's A8 on the same 102 dev games (first 2 items per family; 92 T2-T6). Forcing Qwen3.5-4B to run 16 experiments raised its T2-T6 win rate from 0.000 to 0.191 (paired +0.191 [0.095, 0.287]; P36 supports H1) and narrowed the evidence to a median of 2 consistent classes at the first submission, yet 0.513 of its submissions contradicted evidence it had been shown (A8 on the same games: 0.217; P37 supports H2). Showing the count of consistent rules raised experimenting slightly (+0.42 [0.10, 0.74] a game; P38 supports H3) with no win difference measured (+0.049 [-0.010, 0.107], P39), and once forced, adding the count showed no difference measured either (+0.033 [-0.057, 0.122], P40). By the decision table, the 4B has both a stopping problem and a hypothesis-tracking problem. Descriptively, after forcing an ideal player would win with mean probability 0.70 at the 4B's first submission, above E1's leaders' 0.58-0.62, yet the 4B won 0.29 of the T2-T6 games such a player would be expected to win, against 0.86-0.88 for the leaders (win probabilities over all games including T1, conversion over T2-T6 games; the leaders' figures over E1's games, not these).

By @stw2 via agent · covers #48 · currently selected

Assignment, status and accepted plans remain authoritative on each experiment.

Reported state

Reported progress
E2 completed: three arms run and verified, 14 findings (P28-P41) linked, H1-H3 tested as pre-registered.
Open issues
Why the forced 4B's submissions contradict the evidence: the rule it means against the rule it writes, or no check against the scenes. A10 made 84 wrong submissions with exactly one class consistent, 52 contradicting the evidence and 32 consistent but outside the prior: 0.82 a game, against about 0.3 for E1's leaders over 460 games. One run per arm, 92 headline games, dev only. A8 is E1's model-player run, not a fresh control, and the subset was fixed while its 0 of 92 was known. H1-H3 state their entities as text, so they join the graph through the findings that test them (P36-P38). P1's definition of win probability at the first submission misdescribes diagnose's p_first; a corrected restatement of P1 is being drafted.
Suggested next action
Analyse A10 and A11's wrong submissions made with one class left (contradicting against out-of-prior); test an inference-time consistency check that verifies a candidate rule against the evidence before it is submitted; for training, consider a signal for consistency with the evidence as well as the win.
Access needs
Apple silicon with MLX for further 4B runs; run files held by the owner (restricted).

Exact references

Exact record · P36

3446a2b3-0944-43db-b24d-009be0f3df04

P36: H1's pre-registered test (stopping).

Exact record · P37

7334ea23-17cf-4b45-b0e0-29c77343dad1

P37: H2's pre-registered test (tracking).

Exact record · P38

b571e856-d925-4052-a2d9-f4131060e0fd

P38: H3's pre-registered test (uncertainty).

Exact record · P39

90028119-acaa-45fe-88dc-0edeb4c8ccf7

P39: the count alone, wins against A8.

Exact record · P40

35e725a3-ef63-4c9c-9ff0-2f54bdd21515

P40: the count once forced.

Exact record · P41

01603c0d-e020-457f-8f17-62992068bcfa

P41: forcing when the count is shown.

Exact record · P31

9c6080af-4e48-4962-82be-e9d303ce9238

P31: A8's win rate on E2's games.

Exact record · P32

be062a4b-4e6a-4de7-be80-5e98a9ff5053

P32: A8's behaviour on E2's games.

Exact record · P33

4a45cc3f-4aa0-4f61-8c88-164a0a483b54

P33: the forced arm's behaviour.

Exact record · P28

b4fb1d38-4c66-4640-bfd9-d3f864fa4142

P28: the forced arm's win rate; defines E2's vocabulary.

Experiment · E2

Open experiment →

Accepted plan

Exact plan →