Raises
The intervention raises the outcome against the baseline on the same games: the mean paired difference (arm minus baseline) has a 95% interval entirely above 0.
Hypothesis
Sign in with GitHubHypothesis · H1 · Author-curated prediction
Advisory · cited versions changed
Qwen3.5-4B on ZendoBench 1.0.0 dev, the 102 games of E02's design, one run per arm, inference-time interventions only (no training). Other models, sizes, numbers of forced experiments, splits and the sealed set are out of scope.
finding · premise
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), Qwen3.5-4B (MLX) wins 0.030 (95% Korn-Graubard CI 0.015-0.074); seed-only baseline 0.047. E1 attempt A8.The 4B's T2-T6 win rate without an intervention (3.0 over 430 games): the level any recovery is measured against.
finding · premise
On ZendoBench 1.0.0 dev (460 games), Qwen3.5-4B (MLX) runs 1.91 experiments per game at 0.78 bits of expected information each, and first submits with a median of 98 rule classes still consistent with the evidence (mean win probability 0.09; 460 games with a submission). E1 attempt A8.The 4B's behaviour without an intervention: 1.91 experiments a game, a median of 98 classes alive at its first submission. The interventions act on this behaviour.
finding · premise
Among the seven systems evaluated in E1 on ZendoBench 1.0.0 dev, ranking by experiments per game matches ranking by T2-T6 win rate exactly (Spearman rho 1.00, 7 systems). Descriptive of these systems only; not a causal or population claim.Across E1's seven systems, experiments per game and win rate rank identically. This hypothesis tests one causal reading of that association within one model.
finding · related
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two seed scenes, no experiments) wins 0.047 in expectation.Defines the shared measures used here (win rate T2-T6, experiments per game, classes alive at the first submission) and the seed-only baseline.
Loading research…
The intervention raises the outcome against the baseline on the same games: the mean paired difference (arm minus baseline) has a 95% interval entirely above 0.
Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.
{
"wording": "On the 102 games of E02's design (ZendoBench 1.0.0 dev, 92 of them T2-T6), Qwen3.5-4B forced to run 16 experiments before it may submit wins more T2-T6 games than without an intervention (E1 attempt A8 on the same games): the paired win difference, forced minus A8, has a 95% interval above 0 (ZendoBench compare's pairing and Student-t interval on class means, equal-weight tiers). Read only if forcing narrowed the evidence: the forced arm's median classes alive at the first submission is at most 10.",
"predicate": {
"type": "concept",
"key": "raises"
},
"roles": [
{
"role": "subject",
"definition": "The model and setting under test.",
"value": {
"type": "text",
"value": "Qwen/Qwen3.5-4B bf16, revision 851bf6e8, on ZendoBench 1.0.0's MLX batch driver with E1's settings: Qwen3.5 thinking sampling (temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5), thinking budget 63,487 plus a 2,048-token answer allowance (65,536 a call), 36 games in flight, sampler seed 0, player agent-tools."
}
},
{
"role": "intervention",
"definition": "What changes against the baseline.",
"value": {
"type": "text",
"value": "Forced experimenting: until the player has run 16 experiments, the observation's available_actions is [\"experiment\"] and the answer grammar is ZendoBench's experiment-only variant; the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions)."
}
},
{
"role": "baseline",
"definition": "What the arm is compared with.",
"value": {
"type": "text",
"value": "E1 attempt A8 (Qwen3.5-4B, same model, sampling, hardware and seed, no intervention, player model) on the same 102 games, paired by task ID."
}
},
{
"role": "outcome",
"definition": "The measure compared.",
"value": {
"type": "concept_ref",
"versionId": "7f391bde-3e66-43b6-9a50-d299f7b7e5b2",
"key": "win_rate_t2_t6_equal_weight"
}
},
{
"role": "items",
"definition": "The games the prediction is about.",
"value": {
"type": "text",
"value": "The 102 games of E02's design on ZendoBench 1.0.0 dev: the first 2 items of each family (--per-family 2), T1 10, T2 14, T3 20, T4 10, T5 34, T6 14; 92 T2-T6 games."
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find Learning to test hypotheses, thread Small models: a stopping problem or a hypothesis-tracking problem?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →