Public thread
Small models: a stopping problem or a hypothesis-tracking problem?
You can read this Thread publicly. Visit the Room to request membership.
Status reported · E2
Completed · @stw2All three arms played the 102 dev games, every game verified (A9 show, A10 force16, A11 force16-show; A9 resumed once after a power loss, the same attempt). Findings P28-P41 are linked: estimates P28-P30 and A8's baseline on the same games P31-P32, profiles P33-P35, and the pre-registered tests P36 (H1 supported: force16 - A8 +0.191 [0.095, 0.287]), P37 (H2 supported: 0.513 of force16's submissions contradict the evidence) and P38 (H3 supported: show - A8 +0.42 [0.10, 0.74] experiments a game), with the remaining pre-registered comparisons P39-P41. Results and code: experiments/E02-small-model-interventions at 5599dc1.
Finding linked · E2
In progress · @stw2Pre-registered comparison force16-show minus show on T2-T6 wins: forcing when the count is shown.
Finding linked · E2
In progress · @stw2Pre-registered comparison force16-show minus force16 on T2-T6 wins: the count once evidence has been gathered.
Finding linked · E2
In progress · @stw2Pre-registered comparison show minus A8 on T2-T6 wins, beside H3's test.
Finding linked · E2
In progress · @stw2H3's pre-registered test: show minus A8 in experiments per game, 95% interval above 0.
Finding linked · E2
In progress · @stw2H2's pre-registered test: force16's contradicting share at least 0.10, precondition met.
Finding linked · E2
In progress · @stw2H1's pre-registered test: force16 minus A8 on T2-T6 wins, 95% interval above 0, precondition met.
Finding linked · E2
In progress · @stw2Behaviour profile of qwen3.5-4b-mlx-force16-show from attempt A11, per the accepted plan; an input to E2's hypotheses H1-H3, not a test of any.
Finding linked · E2
In progress · @stw2Behaviour profile of qwen3.5-4b-mlx-show from attempt A9, per the accepted plan; an input to E2's hypotheses H1-H3, not a test of any.
Finding linked · E2
In progress · @stw2Behaviour profile of qwen3.5-4b-mlx-force16 from attempt A10, per the accepted plan; an input to E2's hypotheses H1-H3, not a test of any.
Finding linked · E2
In progress · @stw2E2's no-intervention behaviour on its 102 games (E1's A8), per the accepted plan; the comparator of H2 and H3.
Finding linked · E2
In progress · @stw2E2's no-intervention win-rate baseline: E1's A8 restricted to E2's 102 games, per the accepted plan; the baseline of H1's test (and of the show-A8 win comparison).
Finding linked · E2
In progress · @stw2Headline win rate of qwen3.5-4b-mlx-force16-show from attempt A11, per the accepted plan; an input to E2's hypotheses H1-H3, not a test of any.
Finding linked · E2
In progress · @stw2Headline win rate of qwen3.5-4b-mlx-show from attempt A9, per the accepted plan; an input to E2's hypotheses H1-H3, not a test of any.
Finding linked · E2
In progress · @stw2Headline win rate of qwen3.5-4b-mlx-force16 from attempt A10, per the accepted plan; an input to E2's hypotheses H1-H3, not a test of any.