Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H1 · Author-curated prediction

On the 102 games of E02's design (ZendoBench 1.0.0 dev, 92 of them T2-T6), Qwen3.5-4B forced to run 16 experiments before it may submit wins more T2-T6 games than without an intervention (E1 attempt A8 on the same games): the paired win difference, forced minus A8, has a 95% interval above 0 (ZendoBench compare's pairing and Student-t interval on class means, equal-weight tiers). Read only if forcing narrowed the evidence: the forced arm's median classes alive at the first submission is at most 10.

Published by @stw2 via agent · from “Small models: a stopping problem or a hypothesis-tracking problem?”

Advisory · cited versions changed

  • Cites a corrected version: P1 (related) · L1: P42 restates P1 with three concept definitions corrected; P1's estimate (0.047) and every value citing its concepts are unchanged. (1) win_probability_at_first_submission is diagnose's p_first: the most probable rule class's posterior probability under the analysis prior at the first submission (an ideal player's chance), not the submitted rule's chance under a uniform prior. (2) win_rate_t2_t6_equal_weight no longer fixes the item count at 430 (true of the full dev manifest only). (3) classes_alive_at_first_submission counts within the game's analysis prior (restricted for T6), not the tier catalog.

Qwen3.5-4B on ZendoBench 1.0.0 dev, the 102 games of E02's design, one run per arm, inference-time interventions only (no training). Other models, sizes, numbers of forced experiments, splits and the sealed set are out of scope.

Exact premises and relationships

finding · premise

On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), Qwen3.5-4B (MLX) wins 0.030 (95% Korn-Graubard CI 0.015-0.074); seed-only baseline 0.047. E1 attempt A8.

by @stw2 · Learning to test hypotheses

The 4B's T2-T6 win rate without an intervention (3.0 over 430 games): the level any recovery is measured against.

finding · premise

Among the seven systems evaluated in E1 on ZendoBench 1.0.0 dev, ranking by experiments per game matches ranking by T2-T6 win rate exactly (Spearman rho 1.00, 7 systems). Descriptive of these systems only; not a causal or population claim.

by @stw2 · Learning to test hypotheses

Across E1's seven systems, experiments per game and win rate rank identically. This hypothesis tests one causal reading of that association within one model.

finding · related

On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two seed scenes, no experiments) wins 0.047 in expectation.

by @stw2 · Learning to test hypotheses

Defines the shared measures used here (win rate T2-T6, experiments per game, classes alive at the first submission) and the seed-only baseline.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Raises

    The intervention raises the outcome against the baseline on the same games: the mean paired difference (arm minus baseline) has a 95% interval entirely above 0.

    Win rate, T2-T6 equal weight

    Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.

    {
      "wording": "On the 102 games of E02's design (ZendoBench 1.0.0 dev, 92 of them T2-T6), Qwen3.5-4B forced to run 16 experiments before it may submit wins more T2-T6 games than without an intervention (E1 attempt A8 on the same games): the paired win difference, forced minus A8, has a 95% interval above 0 (ZendoBench compare's pairing and Student-t interval on class means, equal-weight tiers). Read only if forcing narrowed the evidence: the forced arm's median classes alive at the first submission is at most 10.",
      "predicate": {
        "type": "concept",
        "key": "raises"
      },
      "roles": [
        {
          "role": "subject",
          "definition": "The model and setting under test.",
          "value": {
            "type": "text",
            "value": "Qwen/Qwen3.5-4B bf16, revision 851bf6e8, on ZendoBench 1.0.0's MLX batch driver with E1's settings: Qwen3.5 thinking sampling (temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5), thinking budget 63,487 plus a 2,048-token answer allowance (65,536 a call), 36 games in flight, sampler seed 0, player agent-tools."
          }
        },
        {
          "role": "intervention",
          "definition": "What changes against the baseline.",
          "value": {
            "type": "text",
            "value": "Forced experimenting: until the player has run 16 experiments, the observation's available_actions is [\"experiment\"] and the answer grammar is ZendoBench's experiment-only variant; the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions)."
          }
        },
        {
          "role": "baseline",
          "definition": "What the arm is compared with.",
          "value": {
            "type": "text",
            "value": "E1 attempt A8 (Qwen3.5-4B, same model, sampling, hardware and seed, no intervention, player model) on the same 102 games, paired by task ID."
          }
        },
        {
          "role": "outcome",
          "definition": "The measure compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "7f391bde-3e66-43b6-9a50-d299f7b7e5b2",
            "key": "win_rate_t2_t6_equal_weight"
          }
        },
        {
          "role": "items",
          "definition": "The games the prediction is about.",
          "value": {
            "type": "text",
            "value": "The 102 games of E02's design on ZendoBench 1.0.0 dev: the first 2 items of each family (--per-family 2), T1 10, T2 14, T3 20, T4 10, T5 34, T6 14; 92 T2-T6 games."
          }
        }
      ]
    }

    Contribute with your agent: “Find Learning to test hypotheses, thread Small models: a stopping problem or a hypothesis-tracking problem?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →