Flip rate
Share of patched items whose prediction in the target language is correct after the patch; every forward-discordant pair is incorrect in Polish without it.
Hypothesis
Sign in with GitHubHypothesis · H36 · Author-curated prediction
Qwen2.5-1.5B, English source and Polish target, scaffold-final anchor; interpretable only if the validity gate passes (specific flip rate at least 3 times the irrelevant-source rate, breakage at most 5 percentage points) and, after Amendment 1, only at layer pairs whose transport rate is at most 0.35.
finding · premise
On 598 translation-paired Belebele items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.7458 in English and 0.5485 in Polish, an English-minus-Polish gap of 0.1973 (95% interval 0.1558 to 0.2389).The gap whose failures the patch targets.
finding · premise
On 589 translation-paired MMLU items, the likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.5569 in English and 0.3430 in Polish, an English-minus-Polish gap of 0.2139 (95% interval 0.1672 to 0.2606).The gap whose failures the patch targets.
Loading research…
Share of patched items whose prediction in the target language is correct after the patch; every forward-discordant pair is incorrect in Polish without it.
Under the intervention at the anchor and layer pair, the metric on the evaluation set is at least threshold.
Replacement, without scaling, of the residual-stream state entering the target decoder layer at the anchor of the target twin's prompt by the state entering the source decoder layer at the anchor of the other twin's prompt, applied while the target twin's options are scored.
The last token of the prompt: the end of 'Answer (letter):' in English or 'Odpowiedź (litera):' in Polish.
The 326 translation pairs, among 598 Belebele and 589 MMLU pairs left after excluding structurally malformed pairs, that Qwen2.5-1.5B answers correctly in English and incorrectly in Polish by likelihood scoring: 151 Belebele and 175 MMLU pairs.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
{
"wording": "Replacing the Polish twin's residual-stream state at the scaffold-final token with the English twin's state makes Qwen2.5-1.5B answer at least 50% of the forward-discordant translation pairs correctly at the best confirmed pair of source and target layers.",
"predicate": {
"type": "concept",
"key": "rate_at_least"
},
"roles": [
{
"role": "intervention",
"definition": "Change applied while the Polish options are scored.",
"value": {
"type": "concept",
"key": "residual_patch"
}
},
{
"role": "anchor",
"definition": "Token position patched.",
"value": {
"type": "concept",
"key": "scaffold_final_anchor"
}
},
{
"role": "layer_pair",
"definition": "Source and target layers at which the metric is read.",
"value": {
"type": "text",
"value": "the pair with the highest flip rate among pairs confirmed on all forward-discordant pairs"
}
},
{
"role": "model",
"definition": "Model patched.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "evaluation_set",
"definition": "Items patched.",
"value": {
"type": "concept",
"key": "forward_discordant_pairs"
}
},
{
"role": "metric",
"definition": "Quantity bounded from below.",
"value": {
"type": "concept",
"key": "flip_rate"
}
},
{
"role": "threshold",
"definition": "Lower bound on the metric.",
"value": {
"type": "decimal",
"value": "0.50",
"unit": "ratio"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →