Raises
The intervention raises the outcome against the baseline on the same games: the mean paired difference (arm minus baseline) has a 95% interval entirely above 0.
Hypothesis
Sign in with GitHubHypothesis · H3 · Author-curated prediction
Advisory · cited versions changed
Qwen3.5-4B on ZendoBench 1.0.0 dev, the 102 games of E02's design, one run per arm, inference-time interventions only (no training). Other models, sizes, numbers of forced experiments, splits and the sealed set are out of scope.
finding · premise
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), Qwen3.5-4B (MLX) wins 0.030 (95% Korn-Graubard CI 0.015-0.074); seed-only baseline 0.047. E1 attempt A8.The 4B's T2-T6 win rate without an intervention (3.0 over 430 games): the level any recovery is measured against.
finding · premise
On ZendoBench 1.0.0 dev (460 games), Qwen3.5-4B (MLX) runs 1.91 experiments per game at 0.78 bits of expected information each, and first submits with a median of 98 rule classes still consistent with the evidence (mean win probability 0.09; 460 games with a submission). E1 attempt A8.The 4B's behaviour without an intervention: 1.91 experiments a game, a median of 98 classes alive at its first submission. The interventions act on this behaviour.
finding · related
On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two seed scenes, no experiments) wins 0.047 in expectation.Defines the shared measures used here (win rate T2-T6, experiments per game, classes alive at the first submission) and the seed-only baseline.
Loading research…
The intervention raises the outcome against the baseline on the same games: the mean paired difference (arm minus baseline) has a 95% interval entirely above 0.
Metric. Mean number of experiments (scenes the player built and had labelled) per scored game; zendo_bench diagnose field exp_per_game, over all scored dev games including T1.
{
"wording": "On the 102 games of E02's design, Qwen3.5-4B shown after every observation how many rule classes of the game's prior are still consistent with the evidence runs more experiments a game than without an intervention (E1 attempt A8 on the same games): the mean per-game paired difference in experiments, shown minus A8, has a 95% Student-t interval above 0.",
"predicate": {
"type": "concept",
"key": "raises"
},
"roles": [
{
"role": "subject",
"definition": "The model and setting under test.",
"value": {
"type": "text",
"value": "Qwen/Qwen3.5-4B bf16, revision 851bf6e8, on ZendoBench 1.0.0's MLX batch driver with E1's settings: Qwen3.5 thinking sampling (temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5), thinking budget 63,487 plus a 2,048-token answer allowance (65,536 a call), 36 games in flight, sampler seed 0, player agent-tools."
}
},
{
"role": "intervention",
"definition": "What changes against the baseline.",
"value": {
"type": "text",
"value": "The count of consistent rules: after every observation, the observation carries consistent_rules {count, of}, the classes of the game's analysis prior (1,765 for T1-T4, 619 for T5, 427 for T6) consistent with every evidence scene shown, as diagnose counts them, computed from the message alone; the system prompt explains the field."
}
},
{
"role": "baseline",
"definition": "What the arm is compared with.",
"value": {
"type": "text",
"value": "E1 attempt A8 (Qwen3.5-4B, same model, sampling, hardware and seed, no intervention, player model) on the same 102 games, paired by task ID."
}
},
{
"role": "outcome",
"definition": "The measure compared.",
"value": {
"type": "concept_ref",
"versionId": "7f391bde-3e66-43b6-9a50-d299f7b7e5b2",
"key": "experiments_per_game"
}
},
{
"role": "items",
"definition": "The games the prediction is about.",
"value": {
"type": "text",
"value": "The 102 games of E02's design on ZendoBench 1.0.0 dev: the first 2 items of each family (--per-family 2), T1 10, T2 14, T3 20, T4 10, T5 34, T6 14; 92 T2-T6 games."
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find Learning to test hypotheses, thread Small models: a stopping problem or a hypothesis-tracking problem?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →