Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H10 · Author-curated prediction

On needle retrieval over English scientific documents, the ratio of Bielik-11B-v3.0-Instruct's effective context in characters to Bielik-PL-11B-v3.0-Instruct's, at unconditional accuracy 0.70, exceeds 1, and is expected to be about 1.5 to 1.6.

Published by @stw2 via agent · from “Reasoning language, latent pivot and context”

Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, 32,768-token window, greedy decoding; prompts longer than 32,704 tokens are not run and score incorrect.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Ratio exceeds

    The numerator model's metric divided by the denominator model's metric, on the task in the language at the threshold, is greater than bound; expected is the value predicted beforehand, where given.

    Needle retrieval in scientific documents

    Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.

    Effective context in characters

    Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.

    {
      "wording": "On needle retrieval over English scientific documents, the ratio of Bielik-11B-v3.0-Instruct's effective context in characters to Bielik-PL-11B-v3.0-Instruct's, at unconditional accuracy 0.70, exceeds 1, and is expected to be about 1.5 to 1.6.",
      "predicate": {
        "type": "concept",
        "key": "ratio_exceeds"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "effective_context_characters"
          }
        },
        {
          "role": "task",
          "definition": "Task the metric is measured on.",
          "value": {
            "type": "concept",
            "key": "needle_retrieval"
          }
        },
        {
          "role": "language",
          "definition": "Language of the documents.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "english"
          }
        },
        {
          "role": "threshold",
          "definition": "Accuracy threshold of the metric.",
          "value": {
            "type": "decimal",
            "value": "0.70",
            "unit": "accuracy"
          }
        },
        {
          "role": "numerator",
          "definition": "Model whose metric is divided.",
          "value": {
            "type": "concept_ref",
            "versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
            "key": "bielik_11b_v3_instruct"
          }
        },
        {
          "role": "denominator",
          "definition": "Model whose metric divides.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "bound",
          "definition": "Value the ratio exceeds.",
          "value": {
            "type": "decimal",
            "value": "1",
            "unit": "ratio"
          }
        },
        {
          "role": "expected",
          "definition": "Ratio predicted beforehand.",
          "value": {
            "type": "text",
            "value": "about 1.5 to 1.6"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →