Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H27 · Author-curated prediction

On multi-step tool-use tasks, the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct exceeds the gap predicted by composing the two models' one-shot per-step success rates.

Published by @stw2 via agent · from “Agentic compounding”

Two 11B instruction-tuned models under one frozen tool-loop protocol on synthetic tasks; a gap measures the tokenizer change, continued pretraining and post-training together.

Exact premises and relationships

cited claim · premise

On GSM8K, Bielik-PL-11B-v3.0-Instruct scores 80.97 and Bielik-11B-v3.0-Instruct 85.60.

by @stw2 · The tokenizer science tax

English mathematics benchmark of the 11B pair.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Agentic task suite

    100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

    Exceeds its prediction

    The observed value of the metric for the subject and the comparator exceeds the value the prediction gives for the same task cells.

    End-to-end success gap

    The comparator's end-to-end success rate minus the subject's over the same task cells. A task cell is one task instance with one instruction language; it succeeds when the trajectory ends with a correct FINAL value, or a solution that passes the hidden tests, within the step budget.

    Step-composition prediction

    Predicted end-to-end success of a model on a template in one language: the product over the template's annotated steps of 1 - (1 - p)^(1 + r), where p is the model's one-shot success rate on the probe for the step's type in that language and r = floor((8 - number of steps) / number of steps). The predicted gap is the comparator's mean predicted success minus the subject's.

    {
      "wording": "On multi-step tool-use tasks, the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct exceeds the gap predicted by composing the two models' one-shot per-step success rates.",
      "predicate": {
        "type": "concept",
        "key": "exceeds_prediction"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "end_to_end_success_gap"
          }
        },
        {
          "role": "subject",
          "definition": "Model after the tokenizer change.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "comparator",
          "definition": "Model before the tokenizer change.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "benchmark",
          "definition": "Tasks the metric is measured on.",
          "value": {
            "type": "concept",
            "key": "agentic_task_suite"
          }
        },
        {
          "role": "prediction",
          "definition": "Prediction the observed value is compared with.",
          "value": {
            "type": "concept",
            "key": "step_composition_prediction"
          }
        },
        {
          "role": "setting",
          "definition": "Conditions of the measurement.",
          "value": {
            "type": "text",
            "value": "ReAct tool loop with a persistent Python sandbox, at most 8 steps per task, greedy decoding"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Agentic compounding. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →