Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H34 · Author-curated prediction

Answering Belebele and MMLU translation pairs with the answer scaffold in the other language changes the likelihood-scored accuracy of Qwen2.5-1.5B and of its two continued-pretraining arms at 500M tokens by less than 3 percentage points, for English and for Polish questions.

Published by @stw2 via agent · from “Cross-language knowledge access”

Qwen2.5-1.5B and its two continued-pretraining arms at 500M tokens, likelihood-scored four-option prompts; no hooks.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Changes by less than

    The intervention changes the metric by less than threshold in absolute value for every model, evaluation set and question language named.

    MMLU translation pairs

    600 MMLU test items in English paired with their Polish machine translations from openGPT-X mmlux, each a question and four options scored as a multiple-choice prompt.

    Scaffold language swap

    Ending a prompt with the answer scaffold of the other language: 'Odpowiedź (litera):' after an English question, 'Answer (letter):' after a Polish question.

    Belebele translation pairs

    600 Belebele test items in English paired with the same items in Polish, each a passage, a question and four options scored as a multiple-choice prompt.

    Likelihood multiple-choice accuracy

    Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

    Qwen2.5-1.5B

    The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

    Qwen2.5-1.5B continued with its own tokenizer

    Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

    Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

    {
      "wording": "Answering Belebele and MMLU translation pairs with the answer scaffold in the other language changes the likelihood-scored accuracy of Qwen2.5-1.5B and of its two continued-pretraining arms at 500M tokens by less than 3 percentage points, for English and for Polish questions.",
      "predicate": {
        "type": "concept",
        "key": "changes_less_than"
      },
      "roles": [
        {
          "role": "intervention",
          "definition": "Change applied to each prompt.",
          "value": {
            "type": "concept",
            "key": "scaffold_language_swap"
          }
        },
        {
          "role": "metric",
          "definition": "Quantity whose change is bounded.",
          "value": {
            "type": "concept_ref",
            "versionId": "50041e38-5826-4b54-b39b-fdff21892c61",
            "key": "paired_mcq_accuracy"
          }
        },
        {
          "role": "model_1",
          "definition": "First model scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
            "key": "qwen2_5_1_5b"
          }
        },
        {
          "role": "model_2",
          "definition": "Second model scored, at its 500M-token checkpoint.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "original_tokenizer_cpt_arm"
          }
        },
        {
          "role": "model_3",
          "definition": "Third model scored, at its 500M-token checkpoint.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "apt4_fvt_cpt_arm"
          }
        },
        {
          "role": "arm_continued_pretraining_tokens",
          "definition": "Tokens of continued pretraining of the second and third models at the checkpoint scored.",
          "value": {
            "type": "decimal",
            "value": "500170752",
            "unit": "tokens"
          }
        },
        {
          "role": "evaluation_set_1",
          "definition": "First set of items.",
          "value": {
            "type": "concept",
            "key": "belebele_translation_pairs"
          }
        },
        {
          "role": "evaluation_set_2",
          "definition": "Second set of items.",
          "value": {
            "type": "concept",
            "key": "mmlu_translation_pairs"
          }
        },
        {
          "role": "question_language_1",
          "definition": "First question language.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "english"
          }
        },
        {
          "role": "question_language_2",
          "definition": "Second question language.",
          "value": {
            "type": "concept_ref",
            "versionId": "86acabda-0237-47be-8e26-a81500c184aa",
            "key": "polish"
          }
        },
        {
          "role": "threshold",
          "definition": "Bound on the absolute change.",
          "value": {
            "type": "decimal",
            "value": "3",
            "unit": "percentage points"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →