Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H22 · Author-curated prediction

On Polish STEM questions with Polish chain-of-thought, the mean pivot excess of Bielik-PL-11B-v3.0-Instruct is greater than zero.

Published by @stw2 via agent · from “Reasoning language, latent pivot and context”

Teacher-forced logit lens of Bielik-PL-11B-v3.0-Instruct over its existing greedy answers; no new generations.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Greater than

    The metric of the subject on the evaluation_items under the condition, averaged over items where it is defined per item, is greater than the threshold; intervention, where given, is the change the subject went through, and predictor, where given, is the per-item quantity the metric relates to answer correctness.

    Pivot excess

    The band English share of an answer minus its logit-lens English share at layer 50.

    Band English share

    The logit-lens English share averaged over layers 26 to 43 of a 50-layer model, for one answer.

    Logit-lens English share

    At one layer of a model, for the positions of an answer fed back to the model with teacher forcing whose next token is labelled English or Polish by the vocabulary language partition, the probability mass that the model's final normalisation and output head assign to English-labelled pieces when applied to that layer's residual stream, divided by the mass on English- or Polish-labelled pieces, pooled over the positions before the final-answer marker.

    Vocabulary language partition

    A labelling of each vocabulary piece of one tokenizer as English, Polish or other: English when the piece occurs at least 5 times in English corpora at a per-million rate at least 10 times its rate in Polish corpora, Polish by the symmetric rule or when it contains a Polish diacritic, other otherwise (digits, punctuation, bytes, special tokens, ambiguous pieces).

    Polish chain-of-thought

    Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

    Polish STEM question set

    800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.

    APT4 transplant

    Replacement of a Bielik v3 model's Mistral-derived tokenizer with APT4, followed by vocabulary adaptation and post-training.

    {
      "wording": "On Polish STEM questions with Polish chain-of-thought, the mean pivot excess of Bielik-PL-11B-v3.0-Instruct is greater than zero.",
      "predicate": {
        "type": "concept",
        "key": "greater_than"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity predicted to exceed the threshold.",
          "value": {
            "type": "concept",
            "key": "pivot_excess"
          }
        },
        {
          "role": "subject",
          "definition": "Model measured.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "evaluation_items",
          "definition": "Questions whose answers are measured.",
          "value": {
            "type": "concept_ref",
            "versionId": "a8a88632-e5bd-42b6-8a04-17b77ce87d13",
            "key": "polish_stem_questions"
          }
        },
        {
          "role": "condition",
          "definition": "Prompting condition of the answers.",
          "value": {
            "type": "concept_ref",
            "versionId": "d0b79462-c553-4a1f-9612-98ca6d4a0645",
            "key": "polish_chain_of_thought"
          }
        },
        {
          "role": "threshold",
          "definition": "Value the mean is predicted to exceed.",
          "value": {
            "type": "decimal",
            "value": "0",
            "unit": "share"
          }
        },
        {
          "role": "intervention",
          "definition": "Change the model went through.",
          "value": {
            "type": "concept_ref",
            "versionId": "d5bb75c0-40ed-4d4c-afa4-0fded6fbcb94",
            "key": "apt4_transplant"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →