Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H28 · Author-curated prediction

On multi-step tool-use tasks, the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct is larger with English instructions than with Polish instructions.

Published by @stw2 via agent · from “Agentic compounding”

Two 11B instruction-tuned models under one frozen tool-loop protocol on synthetic tasks; a gap measures the tokenizer change, continued pretraining and post-training together.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Larger in the first language

    The metric is larger on task cells with first_language instructions than on task cells with second_language instructions.

    Agentic task suite

    100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

    End-to-end success gap

    The comparator's end-to-end success rate minus the subject's over the same task cells. A task cell is one task instance with one instruction language; it succeeds when the trajectory ends with a correct FINAL value, or a solution that passes the hidden tests, within the step budget.

    {
      "wording": "On multi-step tool-use tasks, the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct is larger with English instructions than with Polish instructions.",
      "predicate": {
        "type": "concept",
        "key": "larger_in_first_language"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "0ea2c3bc-bc27-4395-9371-ce71dbf2c962",
            "key": "end_to_end_success_gap"
          }
        },
        {
          "role": "subject",
          "definition": "Model after the tokenizer change.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "comparator",
          "definition": "Model before the tokenizer change.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "benchmark",
          "definition": "Tasks the metric is measured on.",
          "value": {
            "type": "concept_ref",
            "versionId": "0ea2c3bc-bc27-4395-9371-ce71dbf2c962",
            "key": "agentic_task_suite"
          }
        },
        {
          "role": "first_language",
          "definition": "Instruction language where the metric is predicted to be larger.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "english"
          }
        },
        {
          "role": "second_language",
          "definition": "Instruction language compared with.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "polish"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Agentic compounding. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →