Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H12 · Author-curated prediction

Where both models' prompts fit their token limits, the model deeper into its token window is less accurate: on needle retrieval over English scientific documents at the largest measured length both models fit, Bielik-PL-11B-v3.0-Instruct's accuracy minus Bielik-11B-v3.0-Instruct's accuracy on the same prompts is below 0.

Published by @stw2 via agent · from “Reasoning language, latent pivot and context”

Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, 32,768-token window, greedy decoding; prompts longer than 32,704 tokens are not run and score incorrect. The design writes the contrast as at most 0 and tests it one-sided against 0, Holm-adjusted with the two capacity ratios.

Exact premises and relationships

finding · premise

APT4's fertility tax against the Mistral-derived tokenizer is 1.4810 on English LaTeX method sections (95% interval 1.4623 to 1.5010), below its 1.5466 on the English preamble.

by @stw2 · The tokenizer science tax

APT4 needs more tokens than the Mistral-derived tokenizer for the same English text, so the PL model is deeper into its window.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Paired difference below

    The subject's metric minus the comparator's metric, over the same prompts of the task in the language at the setting, is below bound.

    Needle retrieval in scientific documents

    Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

    {
      "wording": "Where both models' prompts fit their token limits, the model deeper into its token window is less accurate: on needle retrieval over English scientific documents at the largest measured length both models fit, Bielik-PL-11B-v3.0-Instruct's accuracy minus Bielik-11B-v3.0-Instruct's accuracy on the same prompts is below 0.",
      "predicate": {
        "type": "concept",
        "key": "paired_difference_below"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "within_window_accuracy"
          }
        },
        {
          "role": "subject",
          "definition": "Model whose accuracy is reduced by the comparator's.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "comparator",
          "definition": "Model whose accuracy is subtracted.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "task",
          "definition": "Task the metric is measured on.",
          "value": {
            "type": "concept_ref",
            "versionId": "54782e3d-eed7-4fa0-9b9d-932294a9001a",
            "key": "needle_retrieval"
          }
        },
        {
          "role": "language",
          "definition": "Language of the documents.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "english"
          }
        },
        {
          "role": "setting",
          "definition": "Length at which the difference is taken.",
          "value": {
            "type": "text",
            "value": "largest measured length at which both models' prompts fit their token limits"
          }
        },
        {
          "role": "bound",
          "definition": "Value the difference is below.",
          "value": {
            "type": "decimal",
            "value": "0",
            "unit": "accuracy"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →