At least
The outcome, measured on the arm, is at or above the threshold given.
Hypothesis
Sign in with GitHubHypothesis · H2 · Author-curated prediction
Advisory · cited versions changed
Qwen3.5-4B on ZendoBench 1.0.0 dev, the 102 games of E02's design, one run per arm, inference-time interventions only (no training). Other models, sizes, numbers of forced experiments, splits and the sealed set are out of scope.
finding · premise
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), Qwen3.5-4B (MLX) wins 0.030 (95% Korn-Graubard CI 0.015-0.074); seed-only baseline 0.047. E1 attempt A8.The 4B's T2-T6 win rate without an intervention (3.0 over 430 games).
finding · premise
On ZendoBench 1.0.0 dev (460 games), Qwen3.5-4B (MLX) runs 1.91 experiments per game at 0.78 bits of expected information each, and first submits with a median of 98 rule classes still consistent with the evidence (mean win probability 0.09; 460 games with a submission). E1 attempt A8.The 4B's behaviour without an intervention (E1 attempt A8): 1.91 experiments a game, a median of 98 classes alive at its first submission. E1's diagnose of the same attempt also counts 177 of its 915 submissions (19%) contradicting evidence it had seen. The question is whether that persists once forcing has gathered evidence.
finding · related
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two seed scenes, no experiments) wins 0.047 in expectation.Defines the shared measures used here (win rate T2-T6, experiments per game, classes alive at the first submission) and the seed-only baseline.
Loading research…
The outcome, measured on the arm, is at or above the threshold given.
{
"wording": "On the 102 games of E02's design, Qwen3.5-4B forced to run 16 experiments before it may submit still submits rules that contradict the evidence it has seen: at least 10% of its submissions contradict at least one seed, experiment or counterexample shown before them (diagnose's contradicting). Every E1 system that ran 10 or more experiments a game stayed below 10% (GPT-6 Luna 8.7%, DeepSeek-V4-Flash 1.7%, Qwen3.8-27B 1.2%). Read only if forcing narrowed the evidence: the forced arm's median classes alive at the first submission is at most 10.",
"predicate": {
"type": "concept",
"key": "at_least"
},
"roles": [
{
"role": "subject",
"definition": "The model and setting under test.",
"value": {
"type": "text",
"value": "Qwen/Qwen3.5-4B bf16, revision 851bf6e8, on ZendoBench 1.0.0's MLX batch driver with E1's settings: Qwen3.5 thinking sampling (temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5), thinking budget 63,487 plus a 2,048-token answer allowance (65,536 a call), 36 games in flight, sampler seed 0, player agent-tools."
}
},
{
"role": "intervention",
"definition": "What changes against the baseline.",
"value": {
"type": "text",
"value": "Forced experimenting: until the player has run 16 experiments, the observation's available_actions is [\"experiment\"] and the answer grammar is ZendoBench's experiment-only variant; the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions)."
}
},
{
"role": "outcome",
"definition": "The measure compared with the threshold.",
"value": {
"type": "text",
"value": "The share of the arm's submissions that contradict at least one evidence scene shown before them: zendo_bench diagnose's contradicting, over all 102 games."
}
},
{
"role": "threshold",
"definition": "The level the outcome must reach.",
"value": {
"type": "decimal",
"value": "0.10",
"unit": "share of submissions"
}
},
{
"role": "items",
"definition": "The games the prediction is about.",
"value": {
"type": "text",
"value": "The 102 games of E02's design on ZendoBench 1.0.0 dev: the first 2 items of each family (--per-family 2), T1 10, T2 14, T3 20, T4 10, T5 34, T6 14; 92 T2-T6 games."
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find Learning to test hypotheses, thread Small models: a stopping problem or a hypothesis-tracking problem?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →