Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H46 · Author-curated prediction

Patching the English twin's residual state into a Polish pass of Qwen2.5-1.5B whose options are re-rendered in an independently drawn order makes the Polish pass predict its own gold option, on forward-discordant translation pairs whose two gold letters differ, at a rate at least 10 percentage points above the rate under an irrelevant-source patch, at some source and target layer pair.

Published by @stw2 via agent · from “Cross-language knowledge access”

Qwen2.5-1.5B; pairs from the option-permutation audit; pairs whose two gold letters coincide are reported separately and never pooled; the test is void if fewer than 60% of pairs are stable under permutation or if the rate on coinciding-letter pairs equals the rate on the other pairs.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Own-gold rate

    Share of patched items whose prediction in the target pass is the letter of the target pass's own gold option under its current option order.

    Independently permuted option block

    The target pass's options re-rendered in an order drawn independently of the source twin's order, with the gold letter moved with its option.

    Irrelevant-source patch

    The cross-lingual residual patch with the source state taken from the English pass of a different translation pair.

    Exceeds a control by at least

    At some layer_pair, the model's metric on the subset of the evaluation_set under the intervention, with the target_condition applied, exceeds its metric under the control by at least margin.

    Forward-discordant pairs under permutation-marginalised scoring

    Translation pairs, after structural exclusions, that the model answers correctly in English and incorrectly in Polish when each item's prediction is the option with the highest log-probability per continuation token averaged over the cyclic orders of the options.

    Cross-lingual residual patch

    Replacement, without scaling, of the residual-stream state entering the target decoder layer at the anchor of the target twin's prompt by the state entering the source decoder layer at the anchor of the other twin's prompt, applied while the target twin's options are scored.

    Qwen2.5-1.5B

    The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

    {
      "wording": "Patching the English twin's residual state into a Polish pass of Qwen2.5-1.5B whose options are re-rendered in an independently drawn order makes the Polish pass predict its own gold option, on forward-discordant translation pairs whose two gold letters differ, at a rate at least 10 percentage points above the rate under an irrelevant-source patch, at some source and target layer pair.",
      "predicate": {
        "type": "concept",
        "key": "exceeds_control_by_at_least"
      },
      "roles": [
        {
          "role": "intervention",
          "definition": "Change applied while the Polish options are scored.",
          "value": {
            "type": "concept_ref",
            "versionId": "cd49225b-c235-425c-83bc-e24271555357",
            "key": "residual_patch"
          }
        },
        {
          "role": "target_condition",
          "definition": "Rendering of the target pass.",
          "value": {
            "type": "concept",
            "key": "permuted_option_block"
          }
        },
        {
          "role": "control",
          "definition": "Patch the intervention is compared with.",
          "value": {
            "type": "concept",
            "key": "irrelevant_source_patch"
          }
        },
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "own_gold_rate"
          }
        },
        {
          "role": "model",
          "definition": "Model patched.",
          "value": {
            "type": "concept_ref",
            "versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
            "key": "qwen2_5_1_5b"
          }
        },
        {
          "role": "evaluation_set",
          "definition": "Pairs patched.",
          "value": {
            "type": "concept_ref",
            "versionId": "3e921b79-6353-4934-9191-9793faaa652e",
            "key": "permutation_discordant_pairs"
          }
        },
        {
          "role": "subset",
          "definition": "Pairs the metric is read on.",
          "value": {
            "type": "text",
            "value": "pairs whose English gold letter differs from the Polish pass's gold letter after permutation"
          }
        },
        {
          "role": "layer_pair",
          "definition": "Layers at which the metric is read.",
          "value": {
            "type": "text",
            "value": "any scanned pair of source and target decoder layers"
          }
        },
        {
          "role": "margin",
          "definition": "Lower bound on the difference.",
          "value": {
            "type": "decimal",
            "value": "10",
            "unit": "percentage points"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →