Accepted plan

Sign in with GitHub
← Experiment E12

Immutable accepted plan · prospective

Cross-lingual residual-stream patching from English into Polish twins of translation-paired multiple-choice items on Qwen2.5-1.5B, with behavioural controls and two continued-pretraining arms

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 1, committed on 21 July 2026 after the instrument gates and before any scan output existed. The smoke calibration at the pre-registered primary cell (scaffold-final anchor, source layer 23 into target layer 13) flipped 9 of 10 items on each benchmark, and a 12-item probe, whose script and output were not kept, showed another item's English state making the Polish pass predict that item's gold letter at late source layers. The amendment adds a transport-rate gate at every layer pair, fixes the layer grid to odd layers, flags entity spans that fall in the option block, and fixes the scoring path to one forward pass that matched the committed per-item scores at their stored precision. Not part of the amendment: the confirmation-pair selection later implemented in scripts/04_scan.py (at most three pairs within 0.02 of the best flip rate plus at most three transport-clean pairs) differs from the pre-registered rule of all pairs within 2 percentage points of the best, at most six; it is a dated note in DESIGN.md and report.md.

Public source

Plan

Prediction
Patching the English twin's state into the Polish pass makes at least half of the forward-discordant pairs correct at the best confirmed layer pair with the validity gate passing; effective pairs target layers near 13; the entity span carries at least the scaffold-final effect; continued pretraining leaves the flip rate within 10 percentage points; the reverse direction reaches at most half the forward rate; swapping the scaffold language moves accuracy by less than 3 percentage points.
Protocol
01 derives from the committed per-item scores of Qwen2.5-1.5B, with structurally malformed pairs excluded, the pairs answered correctly in English and incorrectly in Polish, gated to 151 Belebele and 175 MMLU pairs, and the reverse set (33 and 49); it freezes with seed 20260703 a scan subsample of 100 items, a within-benchmark derangement for irrelevant sources, a breakage guard of 100 Polish-correct items and shared entity spans. P0 (02, no hooks) scores the three models with the scaffold language swapped and with the English twin's prompt prepended to the Polish prompt. 03 builds the patch: the English twin's residual-stream state entering a source decoder layer at an anchor replaces, without scaling, the Polish twin's state entering a target layer while its options are scored; its smoke gates are a bitwise self-identity check, a byte-identical 10-item repeat and a calibration of every benchmark and anchor. P1 (04) scans the subsample at the scaffold-final anchor over every second layer as source and target, plus the pairs 23 into 13 and 3 into 13. P2 confirms the top P1 pairs (at most 6, all within 2 percentage points of the best) on all forward-discordant pairs with an irrelevant-source control, the breakage guard, and the anchor ladder: scaffold-final token, question-final token, shared entity span (mean-pooled), a random position and the first token. P3 runs the reverse direction, the two continued-pretraining arms at the P2 pairs, and a STEM slice of P2. 05 computes item bootstrap intervals (2000 resamples, seed 20260703) and exact McNemar tests. Amendment 1: every scan cell also records the transport rate under the irrelevant source; the grid is the odd layers 1 to 27 as source and target (196 pairs, containing 23 into 13 and 3 into 13); options are scored from one forward pass over the prompt and the first option's continuation, checked against the committed per-item scores before any result; entity spans carry a question-stem or option-block flag and the anchor comparison is reported by it.
Dataset
Translation pairs of 600 Belebele test items (English and Polish) and 600 MMLU test items with Polish machine translations from openGPT-X mmlux, with committed per-item scores of Qwen2.5-1.5B and its two continued-pretraining arms; the patch sets exclude structurally malformed pairs.
Split
None: confirmation patches every forward-discordant pair; the coarse scan uses a seeded subsample of 100 (50 per benchmark).
Access needs
The two continued-pretraining checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer reference used for the APT4 arm is gated on Hugging Face.
Configurations
As the pre-registered plan, with layers fixed to the odd decoder layers 1 to 27 as source and target, and entity spans flagged as question stem or option block.
Metric
Flip rate: share of forward-discordant pairs whose patched Polish prediction is correct; irrelevant-source flip rate; breakage on Polish-correct pairs; translation-assist recovery of the English-minus-Polish accuracy gap; accuracy change under a scaffold-language swap.
Seeds
20260703 (PCG64: scan subsample, derangement, breakage guard, bootstrap)
Interpretation rule
Validity gate first: specific flip rate at least 3 times the irrelevant-source flip rate, and breakage at most 5 percentage points on the Polish-correct guard; if it fails, patching claims are void and the experiment is reported as an instrument failure. Then the flip rate R at the best confirmed pair: R at least 0.50 is routing-dominant, R at most 0.15 is representation-bound, which also requires the entity-span anchor to fail, otherwise mixed, reported by benchmark. Translation-assist recovery of the gap on the base model, per benchmark: at least 70% comprehension-dominant, at most 30% deeper than comprehension, otherwise mixed. Scaffold swap: absolute accuracy change below 3 percentage points means the gap tracks question language. Arm A's flip rate within 10 percentage points of the base model's with an interval including zero means continued pretraining left routing unchanged. Reverse flip rate at most half the forward rate means an English-sided store. Fixed pairs 23 into 13 and 3 into 13, read only if routing is dominant or mixed: at least 50% of the best R. STEM against non-STEM flip rate is exploratory, without a bin. Amendment 1: transport rate T is the share of irrelevant-source trials whose prediction equals the source item's gold letter (chance 0.25); a layer pair may support a routing claim only if T is at most 0.35, and the routing classification is read only from such pairs; the other bins are unchanged.
Resources
Apple silicon, 128 GB unified memory; bf16 inference with forward pre-hooks on the MPS backend, about 5 to 7 hours; no training.
Prior work
arXiv:2607.08393 patches residual states between layers of the same failing prompt (self-patching), reports effective pairs from about 0.10L and 0.82L into 0.45L, and uses an irrelevant-patching control.

Selected exact hypotheses and premises

Hypothesis · H34

cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Whether the gap tracks question language rather than scaffold format.

Hypothesis · H35

489055f5-af11-44e3-b63d-c59d6d7cfef6

How much of the gap the English twin in the prompt recovers.

Hypothesis · H36

cd49225b-c235-425c-83bc-e24271555357

The keystone prediction: routing-bound failures reachable by the patch.

Hypothesis · H37

fe0d162e-9992-4466-9b39-c6fb4a6ff108

Where effective layer pairs lie.

Hypothesis · H38

72ecb52b-0fef-4a96-8a5b-4ec81fb1e202

Which token positions carry the effect.

Hypothesis · H39

e73869df-b1f8-4941-873f-5b3ba6931164

Whether continued pretraining changes the patch effect.

Hypothesis · H40

764aaab5-717b-4376-abb9-30c3788c94b1

Whether the effect is asymmetric between directions.

Hypothesis · H41

69a5ef7f-14e9-4856-a165-fffa8dec769e

Whether two fixed layer pairs suffice.

Premise · P244

e6ce775f-1927-4ab2-90fc-cdab7bcec624

The English-minus-Polish accuracy gap of the base model on these translation pairs.

Premise · P247

22e9d053-dd03-4cb2-ba2d-bea9e71a1260

The English-minus-Polish accuracy gap of the base model on these translation pairs.

Premise · P264

72f15189-34bd-468a-a7dd-44743ac5531b

Continued pretraining bought no Polish access gain on the same models.