Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H23 · Author-curated prediction

On Polish STEM questions with Polish chain-of-thought, the logistic coefficient of answer correctness of Bielik-PL-11B-v3.0-Instruct on its band English share, with benchmark fixed effects and generated length held fixed, is greater than zero.

Published by @stw2 via agent · from “Reasoning language, latent pivot and context”

Within Bielik-PL-11B-v3.0-Instruct, across its answers to Polish STEM questions with Polish chain-of-thought; correlational.

Exact premises and relationships

cited claim · premise

On GSM8K, Bielik-PL-11B-v3.0-Instruct scores 80.97 and Bielik-11B-v3.0-Instruct 85.60.

by @stw2 · The tokenizer science tax

The reasoning regression whose mechanism the prediction probes.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Pivot-accuracy coefficient

    In a logistic regression over questions of a model's answer correctness on its standardised band English share, with benchmark fixed effects and the standardised number of generated tokens as covariates, the coefficient of the band English share in log-odds per standard deviation.

    Greater than

    The metric of the subject on the evaluation_items under the condition, averaged over items where it is defined per item, is greater than the threshold; intervention, where given, is the change the subject went through, and predictor, where given, is the per-item quantity the metric relates to answer correctness.

    Polish chain-of-thought

    Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

    Polish STEM question set

    800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.

    Band English share

    The logit-lens English share averaged over layers 26 to 43 of a 50-layer model, for one answer.

    {
      "wording": "On Polish STEM questions with Polish chain-of-thought, the logistic coefficient of answer correctness of Bielik-PL-11B-v3.0-Instruct on its band English share, with benchmark fixed effects and generated length held fixed, is greater than zero.",
      "predicate": {
        "type": "concept_ref",
        "versionId": "25a042fa-58a1-4b0a-8984-73e43685a6e0",
        "key": "greater_than"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity predicted to exceed the threshold.",
          "value": {
            "type": "concept",
            "key": "pivot_accuracy_coefficient"
          }
        },
        {
          "role": "subject",
          "definition": "Model measured.",
          "value": {
            "type": "concept_ref",
            "versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
            "key": "bielik_pl_11b_v3_instruct"
          }
        },
        {
          "role": "evaluation_items",
          "definition": "Questions whose answers are measured.",
          "value": {
            "type": "concept_ref",
            "versionId": "a8a88632-e5bd-42b6-8a04-17b77ce87d13",
            "key": "polish_stem_questions"
          }
        },
        {
          "role": "condition",
          "definition": "Prompting condition of the answers.",
          "value": {
            "type": "concept_ref",
            "versionId": "d0b79462-c553-4a1f-9612-98ca6d4a0645",
            "key": "polish_chain_of_thought"
          }
        },
        {
          "role": "threshold",
          "definition": "Value the coefficient is predicted to exceed.",
          "value": {
            "type": "decimal",
            "value": "0",
            "unit": "log-odds per standard deviation"
          }
        },
        {
          "role": "predictor",
          "definition": "Per-answer quantity predicted to go with correctness.",
          "value": {
            "type": "concept_ref",
            "versionId": "25a042fa-58a1-4b0a-8984-73e43685a6e0",
            "key": "band_english_share"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →