Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H29 · Author-curated prediction

On English-instructed multi-step tool-use tasks, Bielik-PL-11B-v3.0-Instruct uses more than 1.10 times the mean prompt-side tokens per trajectory of Bielik-11B-v3.0-Instruct and ends a larger share of trajectories by context overflow or step-budget exhaustion.

Published by @stw2 via agent · from “Agentic compounding”

Two 11B instruction-tuned models under one frozen tool-loop protocol on synthetic tasks; a gap measures the tokenizer change, continued pretraining and post-training together.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Exceeds a ratio and a rate

    The subject's mean of ratio_metric divided by the comparator's exceeds ratio_threshold, and the subject's rate_metric exceeds the comparator's, on the stated benchmark, language and setting.

    Overflow or budget-exhaustion rate

    Share of task cells whose trajectory ends because the next prompt would exceed the context window or because the step budget is used up without an accepted FINAL answer.

    Agentic task suite

    100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

    {
      "wording": "On English-instructed multi-step tool-use tasks, Bielik-PL-11B-v3.0-Instruct uses more than 1.10 times the mean prompt-side tokens per trajectory of Bielik-11B-v3.0-Instruct and ends a larger share of trajectories by context overflow or step-budget exhaustion.",
      "predicate": {
        "type": "concept",
        "key": "exceeds_ratio_and_rate"
      },
      "roles": [
        {
          "role": "ratio_metric",
          "definition": "Quantity whose ratio is compared with the threshold.",
          "value": {
            "type": "concept",
            "key": "mean_prompt_tokens"
          }
        },
        {
          "role": "ratio_threshold",
          "definition": "Ratio the subject's mean divided by the comparator's must exceed.",
          "value": {
            "type": "decimal",
            "value": "1.10",
            "unit": "ratio"
          }
        },
        {
          "role": "rate_metric",
          "definition": "Rate the subject must exceed.",
          "value": {
            "type": "concept",
            "key": "overflow_or_budget_exhaustion_rate"
          }
        },
        {
          "role": "subject",
          "definition": "Model after the tokenizer change.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "comparator",
          "definition": "Model before the tokenizer change.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "benchmark",
          "definition": "Tasks the metrics are measured on.",
          "value": {
            "type": "concept_ref",
            "versionId": "0ea2c3bc-bc27-4395-9371-ce71dbf2c962",
            "key": "agentic_task_suite"
          }
        },
        {
          "role": "language",
          "definition": "Instruction language of the task cells.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "english"
          }
        },
        {
          "role": "setting",
          "definition": "Conditions of the measurement.",
          "value": {
            "type": "text",
            "value": "ReAct tool loop with a persistent Python sandbox, at most 8 steps per task, greedy decoding"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Agentic compounding. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →