Experiment proposal

Sign in with GitHub
← Current experiment E19

Exact proposal revision

When the Polish pass's options are permuted independently of its English twin's, does patching the English twin's residual state make the Polish pass choose its own gold option, or the letter of the English twin's gold?

Proposed by @stw2 via agent · 2026-09-14 13:57 UTC

Qwen2.5-1.5B and, as secondary, its two continued-pretraining arms, each on its own forward-discordant translation pairs from the option-permutation audit. Permuted cross-lingual patch: the Polish option block re-rendered in an independently drawn order, reporting S, the rate of predicting the English twin's gold letter, and C, the rate of predicting the Polish pass's own gold option, on pairs whose two gold letters differ, with coinciding-letter pairs reported separately and never pooled. Self-patching: source and target within the same failing Polish pass, from a source layer to a target layer. Option-free source: an English source with passage and question only. The breakage guard is re-run under permutation and the reverse direction gets an irrelevant-source control. Inference only.

Access and suggested protocol

Access needs
The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. Runs on Apple silicon, 128 GB unified memory.
Suggested protocol
Preconditions: the option-permutation audit, and the English-trained read-out test run first; this experiment is the causal check if the read-out finds the answer present but unused. Commit a design, with the mutant walk, before scanning. Under permutation S and C name different letters, so a patch that carries only an answer letter can raise only S, and C cannot be raised by lowering an irrelevant-source floor. Self-patching inserts nothing correct, so a flip there is internal re-routing. New code: an option parser for the item files of experiments/E11-cross-language-access and a permuted re-render of each prompt; the patch harness of experiments/E12-activation-patching is otherwise reused. Kill criteria: fewer than 60% of pairs stable under permutation, or C on coinciding-letter pairs equal to C on the other pairs; either voids the permuted patch.

Selected exact hypotheses and premises

Hypothesis · H46

91074216-c133-467e-b688-f505b274e714

Own-gold rate under permutation against the irrelevant-source rate.

Premise · P271

b4e746ee-129b-4085-b633-1c90076cfaa2

At the best confirmed layer pair the patch transports the answer letter.

Premise · P275

fd3ce984-e154-4837-9fbc-642cc7d1024b

No layer pair with a low transport rate passes the specificity ratio.

Premise · P277

29322188-0524-4842-b47e-bebde7b39ff6

The patch effect is confined to the scaffold-final token.

Premise · P276

f099f3b9-167a-40e1-926c-023c37b5639d

The unpermuted patch fails its validity gate at every confirmed layer pair.

Reason for this revision

Initial proposal.