Own-gold rate
Share of patched items whose prediction in the target pass is the letter of the target pass's own gold option under its current option order.
Hypothesis
Sign in with GitHubHypothesis · H46 · Author-curated prediction
Qwen2.5-1.5B; pairs from the option-permutation audit; pairs whose two gold letters coincide are reported separately and never pooled; the test is void if fewer than 60% of pairs are stable under permutation or if the rate on coinciding-letter pairs equals the rate on the other pairs.
finding · premise
At source layer 25 and target layer 25 of Qwen2.5-1.5B, the scaffold-final residual patch has flip rate 0.9663, irrelevant-source flip rate 0.2362 and transport rate 0.9632 on 326 forward-discordant pairs, and breakage 0.1600 on 100 Polish-correct pairs; the transport rate is above 0.35.At the best confirmed layer pair the patch transports the answer letter.
finding · premise
For the scaffold-final residual patch on 100 forward-discordant pairs of Qwen2.5-1.5B, 0 of the 152 odd-layer source and target pairs, among 196, whose transport rate is at most 0.35 have a flip rate at least 3 times their irrelevant-source flip rate; the highest ratio is 1.7500, the highest flip rate 0.3800, and the mean flip rate 0.1747 against a mean irrelevant-source flip rate of 0.1661.No layer pair with a low transport rate passes the specificity ratio.
Loading research…
Share of patched items whose prediction in the target pass is the letter of the target pass's own gold option under its current option order.
The target pass's options re-rendered in an order drawn independently of the source twin's order, with the gold letter moved with its option.
The cross-lingual residual patch with the source state taken from the English pass of a different translation pair.
At some layer_pair, the model's metric on the subset of the evaluation_set under the intervention, with the target_condition applied, exceeds its metric under the control by at least margin.
Translation pairs, after structural exclusions, that the model answers correctly in English and incorrectly in Polish when each item's prediction is the option with the highest log-probability per continuation token averaged over the cyclic orders of the options.
Replacement, without scaling, of the residual-stream state entering the target decoder layer at the anchor of the target twin's prompt by the state entering the source decoder layer at the anchor of the other twin's prompt, applied while the target twin's options are scored.
The 1.5B-parameter base language model of the Qwen2.5 series, without further training.
{
"wording": "Patching the English twin's residual state into a Polish pass of Qwen2.5-1.5B whose options are re-rendered in an independently drawn order makes the Polish pass predict its own gold option, on forward-discordant translation pairs whose two gold letters differ, at a rate at least 10 percentage points above the rate under an irrelevant-source patch, at some source and target layer pair.",
"predicate": {
"type": "concept",
"key": "exceeds_control_by_at_least"
},
"roles": [
{
"role": "intervention",
"definition": "Change applied while the Polish options are scored.",
"value": {
"type": "concept_ref",
"versionId": "cd49225b-c235-425c-83bc-e24271555357",
"key": "residual_patch"
}
},
{
"role": "target_condition",
"definition": "Rendering of the target pass.",
"value": {
"type": "concept",
"key": "permuted_option_block"
}
},
{
"role": "control",
"definition": "Patch the intervention is compared with.",
"value": {
"type": "concept",
"key": "irrelevant_source_patch"
}
},
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "own_gold_rate"
}
},
{
"role": "model",
"definition": "Model patched.",
"value": {
"type": "concept_ref",
"versionId": "b7e7c5ce-2805-43d7-9f89-ccda038df1b5",
"key": "qwen2_5_1_5b"
}
},
{
"role": "evaluation_set",
"definition": "Pairs patched.",
"value": {
"type": "concept_ref",
"versionId": "3e921b79-6353-4934-9191-9793faaa652e",
"key": "permutation_discordant_pairs"
}
},
{
"role": "subset",
"definition": "Pairs the metric is read on.",
"value": {
"type": "text",
"value": "pairs whose English gold letter differs from the Polish pass's gold letter after permutation"
}
},
{
"role": "layer_pair",
"definition": "Layers at which the metric is read.",
"value": {
"type": "text",
"value": "any scanned pair of source and target decoder layers"
}
},
{
"role": "margin",
"definition": "Lower bound on the difference.",
"value": {
"type": "decimal",
"value": "10",
"unit": "percentage points"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →