Attributed research checkpoint
E2 is complete: A9-A11 against E1's A8 on the same 102 dev games (first 2 items per family; 92 T2-T6). Forcing Qwen3.5-4B to run 16 experiments raised its T2-T6 win rate from 0.000 to 0.191 (paired +0.191 [0.095, 0.287]; P36 supports H1) and narrowed the evidence to a median of 2 consistent classes at the first submission, yet 0.513 of its submissions contradicted evidence it had been shown (A8 on the same games: 0.217; P37 supports H2). Showing the count of consistent rules raised experimenting slightly (+0.42 [0.10, 0.74] a game; P38 supports H3) with no win difference measured (+0.049 [-0.010, 0.107], P39), and once forced, adding the count showed no difference measured either (+0.033 [-0.057, 0.122], P40). By the decision table, the 4B has both a stopping problem and a hypothesis-tracking problem. Descriptively, after forcing an ideal player would win with mean probability 0.70 at the 4B's first submission, above E1's leaders' 0.58-0.62, yet the 4B won 0.29 of the T2-T6 games such a player would be expected to win, against 0.86-0.88 for the leaders (win probabilities over all games including T1, conversion over T2-T6 games; the leaders' figures over E1's games, not these).
Assignment, status and accepted plans remain authoritative on each experiment.
Reported state
- Reported progress
- E2 completed: three arms run and verified, 14 findings (P28-P41) linked, H1-H3 tested as pre-registered.
- Open issues
- Why the forced 4B's submissions contradict the evidence: the rule it means against the rule it writes, or no check against the scenes. A10 made 84 wrong submissions with exactly one class consistent, 52 contradicting the evidence and 32 consistent but outside the prior: 0.82 a game, against about 0.3 for E1's leaders over 460 games. One run per arm, 92 headline games, dev only. A8 is E1's model-player run, not a fresh control, and the subset was fixed while its 0 of 92 was known. H1-H3 state their entities as text, so they join the graph through the findings that test them (P36-P38). P1's definition of win probability at the first submission misdescribes diagnose's p_first; a corrected restatement of P1 is being drafted.
- Suggested next action
- Analyse A10 and A11's wrong submissions made with one class left (contradicting against out-of-prior); test an inference-time consistency check that verifies a candidate rule against the evidence before it is submitted; for training, consider a signal for consistency with the evidence as well as the win.
- Access needs
- Apple silicon with MLX for further 4B runs; run files held by the owner (restricted).
Exact references
Exact record · P28
b4fb1d38-4c66-4640-bfd9-d3f864fa4142P28: the forced arm's win rate; defines E2's vocabulary.
Experiment · E2
Open experiment →Accepted plan
Exact plan →