Public research
Publications
Cited claims and findings from everyone in Learning to test hypotheses. Each row opens the exact published assertion and its provenance.
Wording
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two seed scenes, no experiments) wins 0.047 in expectation.e6d5a46f-37c8-4977-9314-30756c30fa6a · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), when Qwen3.5-4B (MLX) is shown how many rule classes are still consistent, also forcing it to run 16 experiments before it may submit raises its T2-T6 win rate by 0.175: paired difference, 95% CI 0.072-0.278. E2 attempt A11 against A9.01603c0d-e020-457f-8f17-62992068bcfa · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), once Qwen3.5-4B (MLX) is forced to run 16 experiments before it may submit, the T2-T6 win rate with the count of consistent rule classes also shown differs by +0.033 from the win rate without it: paired difference, 95% CI -0.057 to 0.122; no difference measured. E2 attempt A11 against A10.35e725a3-ef63-4c9c-9ff0-2f54bdd21515 · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) shown after every observation how many rule classes are still consistent has a T2-T6 win rate that differs by +0.049 from the same model's without an intervention (E1's A8 on the same games): paired difference, 95% CI -0.010 to 0.107; no difference measured. E2 attempt A9 against E1 attempt A8.90028119-acaa-45fe-88dc-0edeb4c8ccf7 · author-curated
On the 102 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family), Qwen3.5-4B (MLX) shown after every observation how many rule classes are still consistent with the evidence runs 0.42 more experiments per game than without an intervention (E1's A8 on the same games): paired per-game difference, 95% CI 0.10-0.74; 2.30 against 1.88. E2 attempt A9 against E1 attempt A8.b571e856-d925-4052-a2d9-f4131060e0fd · author-curated
On the 102 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit contradicts evidence it has already been shown in 97 of its 189 submissions (0.513), above H2's pre-registered threshold of 0.10; without an intervention (E1's A8 on the same games) the share is 0.217 (44 of 203). E2 attempt A10.7334ea23-17cf-4b45-b0e0-29c77343dad1 · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit has a T2-T6 win rate 0.191 higher than without an intervention (E1's A8 on the same games): paired difference, 95% CI 0.095-0.287. The forced arm first submitted with a median of 2 consistent rule classes, so the precondition for reading H1 holds. E2 attempt A10 against E1 attempt A8.3446a2b3-0944-43db-b24d-009be0f3df04 · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) without an intervention, E1's run restricted to these games, has a T2-T6 win rate of 0.000 (95% Korn-Graubard CI 0.000-0.172); its 12 T2-T6 wins over all 430 headline dev games fall outside these games. E1 attempt A8.9c6080af-4e48-4962-82be-e9d303ce9238 · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit and shown after every observation how many rule classes are still consistent has a T2-T6 win rate of 0.224 (95% Korn-Graubard CI 0.136-0.387); the seed-only baseline on the same games is 0.033. E2 attempt A11.43fec666-c5a7-4ca2-8213-93e3110bfa10 · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) shown after every observation how many rule classes are still consistent with the evidence has a T2-T6 win rate of 0.049 (95% Korn-Graubard CI 0.017-0.215); the seed-only baseline on the same games is 0.033. E2 attempt A9.8cfab1fc-54aa-4522-8294-e0ec5f06fe4a · author-curated
On the 92 T2-T6 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family; equal tier weights), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit has a T2-T6 win rate of 0.191 (95% Korn-Graubard CI 0.114-0.348); the seed-only baseline on the same games is 0.033. E2 attempt A10.b4fb1d38-4c66-4640-bfd9-d3f864fa4142 · author-curated
On ZendoBench 1.0.0 dev (T2-T6, 430 games, paired, equal tier weights), GPT-6 Luna at reasoning effort high wins 0.087 more (95% CI 0.044-0.130) through OpenRouter (E1 attempt A2) than through the Codex CLI with no tools (E3 attempt A12); by tier, in points: T2 +0.0, T3 +8.0, T4 +3.3, T5 +19.0, T6 +13.3. Every Codex call returned a complete answer; which difference between the harnesses causes the gap is not identified.f4ff537a-9cc4-4b5c-8550-ed3022a83291 · author-curated
On ZendoBench 1.0.0 dev (460 games), GPT-6 Luna (reasoning high, Codex CLI, no tools) runs 15.38 experiments per game at 0.42 bits of expected information each, and first submits with a median of 2 rule classes still consistent with the evidence (mean win probability 0.58; 421 games with a submission). E3 attempt A12.80a9d061-2941-46bc-b9b5-5596f5ab324b · author-curated
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), GPT-6 Luna (reasoning high, Codex CLI, no tools) wins 0.362 (95% Korn-Graubard CI 0.320-0.411); seed-only baseline 0.047. E3 attempt A12.37b35a96-4080-4119-afc4-aedf1144cc63 · author-curated
On ZendoBench 1.0.0 dev (460 games), GPT-5.6 Terra (reasoning high, Codex CLI, no tools) runs 11.08 experiments per game at 0.52 bits of expected information each, and first submits with a median of 3.5 rule classes still consistent with the evidence (mean win probability 0.45; 460 games with a submission). E3 attempt A15.73127491-c89e-4685-bcad-48a9ff5e5cf4 · author-curated
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), GPT-5.6 Terra (reasoning high, Codex CLI, no tools) wins 0.408 (95% Korn-Graubard CI 0.369-0.451); seed-only baseline 0.047. E3 attempt A15.602a6b09-0683-4dc2-aebd-d042736f2922 · author-curated
On ZendoBench 1.0.0 dev (460 games), GPT-6 Sol (reasoning high, Codex CLI, no tools) runs 19.23 experiments per game at 0.34 bits of expected information each, and first submits with a median of 2 rule classes still consistent with the evidence (mean win probability 0.68; 460 games with a submission). E3 attempt A13.afbcc984-197a-4e2c-89ab-3150d42e6c69 · author-curated
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), GPT-6 Sol (reasoning high, Codex CLI, no tools) wins 0.692 (95% Korn-Graubard CI 0.646-0.735); seed-only baseline 0.047. E3 attempt A13.83d200cc-3c36-4cde-8eae-3c1ad5268afb · author-curated
On ZendoBench 1.0.0 dev (460 games), GPT-6.1 Sol (reasoning high, Codex CLI, no tools) runs 18.10 experiments per game at 0.37 bits of expected information each, and first submits with a median of 1 rule class still consistent with the evidence (mean win probability 0.73; 460 games with a submission). E3 attempt A14.f11a20c9-8758-461d-a3cb-624ca6b2500b · author-curated
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), GPT-6.1 Sol (reasoning high, Codex CLI, no tools) wins 0.897 (95% Korn-Graubard CI 0.847-0.927); seed-only baseline 0.047. E3 attempt A14.8117a6b7-ce4c-46c7-9cfc-1b2fe62afcd9 · author-curated
On ZendoBench 1.0.0 dev (460 games), GPT-6 Astra (reasoning high, Codex CLI, no tools) runs 17.14 experiments per game at 0.39 bits of expected information each, and first submits with a median of 1 rule class still consistent with the evidence (mean win probability 0.70; 460 games with a submission). E3 attempt A16.1b1d403d-5ca4-47a0-bf4d-b020e57fdb0a · author-curated
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), GPT-6 Astra (reasoning high, Codex CLI, no tools) wins 0.943 (95% Korn-Graubard CI 0.903-0.963); seed-only baseline 0.047. E3 attempt A16.3143857e-1e97-4667-88f1-16c8228d167f · author-curated
Among the seven systems evaluated in E1 on ZendoBench 1.0.0 dev, ranking by experiments per game matches ranking by T2-T6 win rate exactly (Spearman rho 1.00, 7 systems). Descriptive of these systems only; not a causal or population claim.da8d956f-13f3-4dbc-b0e0-bb886394b401 · author-curated
On ZendoBench 1.0.0 dev (460 games), Qwen3.5-2B (MLX) runs 0.21 experiments per game at 0.35 bits of expected information each, and first submits with a median of 188 rule classes still consistent with the evidence (mean win probability 0.06; 460 games with a submission). E1 attempt A1.cd29dc79-e65d-4baa-a036-42920fbc748b · author-curated
On ZendoBench 1.0.0 dev (460 games), Qwen3.5-4B (MLX) runs 1.91 experiments per game at 0.78 bits of expected information each, and first submits with a median of 98 rule classes still consistent with the evidence (mean win probability 0.09; 460 games with a submission). E1 attempt A8.53f39276-69b9-4260-aa09-af060ba7c2ee · author-curated
On ZendoBench 1.0.0 dev (460 games), Qwen3.5-27B-FP8 (vLLM, H100) runs 5.36 experiments per game at 0.70 bits of expected information each, and first submits with a median of 27 rule classes still consistent with the evidence (mean win probability 0.19; 404 games with a submission). E1 attempt A6.f164a7c1-d061-4a5f-83ea-e6980601af35 · author-curated
30 loaded