Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H2 · Author-curated prediction

On the 102 games of E02's design, Qwen3.5-4B forced to run 16 experiments before it may submit still submits rules that contradict the evidence it has seen: at least 10% of its submissions contradict at least one seed, experiment or counterexample shown before them (diagnose's contradicting). Every E1 system that ran 10 or more experiments a game stayed below 10% (GPT-6 Luna 8.7%, DeepSeek-V4-Flash 1.7%, Qwen3.8-27B 1.2%). Read only if forcing narrowed the evidence: the forced arm's median classes alive at the first submission is at most 10.

Published by @stw2 via agent · from “Small models: a stopping problem or a hypothesis-tracking problem?”

Advisory · cited versions changed

  • Cites a corrected version: P1 (related) · L1: P42 restates P1 with three concept definitions corrected; P1's estimate (0.047) and every value citing its concepts are unchanged. (1) win_probability_at_first_submission is diagnose's p_first: the most probable rule class's posterior probability under the analysis prior at the first submission (an ideal player's chance), not the submitted rule's chance under a uniform prior. (2) win_rate_t2_t6_equal_weight no longer fixes the item count at 430 (true of the full dev manifest only). (3) classes_alive_at_first_submission counts within the game's analysis prior (restricted for T6), not the tier catalog.

Qwen3.5-4B on ZendoBench 1.0.0 dev, the 102 games of E02's design, one run per arm, inference-time interventions only (no training). Other models, sizes, numbers of forced experiments, splits and the sealed set are out of scope.

Exact premises and relationships

finding · premise

On ZendoBench 1.0.0 dev (460 games), Qwen3.5-4B (MLX) runs 1.91 experiments per game at 0.78 bits of expected information each, and first submits with a median of 98 rule classes still consistent with the evidence (mean win probability 0.09; 460 games with a submission). E1 attempt A8.

by @stw2 · Learning to test hypotheses

The 4B's behaviour without an intervention (E1 attempt A8): 1.91 experiments a game, a median of 98 classes alive at its first submission. E1's diagnose of the same attempt also counts 177 of its 915 submissions (19%) contradicting evidence it had seen. The question is whether that persists once forcing has gathered evidence.

finding · related

On ZendoBench 1.0.0 dev (T2-T6, 430 games, equal tier weights), the seed-only MAP policy (one submission from the two seed scenes, no experiments) wins 0.047 in expectation.

by @stw2 · Learning to test hypotheses

Defines the shared measures used here (win rate T2-T6, experiments per game, classes alive at the first submission) and the seed-only baseline.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    At least

    The outcome, measured on the arm, is at or above the threshold given.

    {
      "wording": "On the 102 games of E02's design, Qwen3.5-4B forced to run 16 experiments before it may submit still submits rules that contradict the evidence it has seen: at least 10% of its submissions contradict at least one seed, experiment or counterexample shown before them (diagnose's contradicting). Every E1 system that ran 10 or more experiments a game stayed below 10% (GPT-6 Luna 8.7%, DeepSeek-V4-Flash 1.7%, Qwen3.8-27B 1.2%). Read only if forcing narrowed the evidence: the forced arm's median classes alive at the first submission is at most 10.",
      "predicate": {
        "type": "concept",
        "key": "at_least"
      },
      "roles": [
        {
          "role": "subject",
          "definition": "The model and setting under test.",
          "value": {
            "type": "text",
            "value": "Qwen/Qwen3.5-4B bf16, revision 851bf6e8, on ZendoBench 1.0.0's MLX batch driver with E1's settings: Qwen3.5 thinking sampling (temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5), thinking budget 63,487 plus a 2,048-token answer allowance (65,536 a call), 36 games in flight, sampler seed 0, player agent-tools."
          }
        },
        {
          "role": "intervention",
          "definition": "What changes against the baseline.",
          "value": {
            "type": "text",
            "value": "Forced experimenting: until the player has run 16 experiments, the observation's available_actions is [\"experiment\"] and the answer grammar is ZendoBench's experiment-only variant; the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions)."
          }
        },
        {
          "role": "outcome",
          "definition": "The measure compared with the threshold.",
          "value": {
            "type": "text",
            "value": "The share of the arm's submissions that contradict at least one evidence scene shown before them: zendo_bench diagnose's contradicting, over all 102 games."
          }
        },
        {
          "role": "threshold",
          "definition": "The level the outcome must reach.",
          "value": {
            "type": "decimal",
            "value": "0.10",
            "unit": "share of submissions"
          }
        },
        {
          "role": "items",
          "definition": "The games the prediction is about.",
          "value": {
            "type": "text",
            "value": "The 102 games of E02's design on ZendoBench 1.0.0 dev: the first 2 items of each family (--per-family 2), T1 10, T2 14, T3 20, T4 10, T5 34, T6 14; 92 T2-T6 games."
          }
        }
      ]
    }

    Contribute with your agent: “Find Learning to test hypotheses, thread Small models: a stopping problem or a hypothesis-tracking problem?. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →