Accepted plan

Sign in with GitHub
← Experiment E12

Immutable accepted plan · prospective

Cross-lingual residual-stream patching from English into Polish twins of translation-paired multiple-choice items on Qwen2.5-1.5B, with behavioural controls and two continued-pretraining arms

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

The pre-registered design, locked on 20 July 2026 before any measurement; its commit holds the design only.

Public source

https://github.com/stw2/tokenizer-science-tax @ 6819d364ee1992d0734a3cea7f5e54de0cb8e74e

Reference checked 2026-09-14 13:18 UTC. No code was executed or scientific result verified.

Plan

Prediction
Patching the English twin's state into the Polish pass makes at least half of the forward-discordant pairs correct at the best confirmed layer pair with the validity gate passing; effective pairs target layers near 13; the entity span carries at least the scaffold-final effect; continued pretraining leaves the flip rate within 10 percentage points; the reverse direction reaches at most half the forward rate; swapping the scaffold language moves accuracy by less than 3 percentage points.
Protocol
01 derives from the committed per-item scores of Qwen2.5-1.5B, with structurally malformed pairs excluded, the pairs answered correctly in English and incorrectly in Polish, gated to 151 Belebele and 175 MMLU pairs, and the reverse set (33 and 49); it freezes with seed 20260703 a scan subsample of 100 items, a within-benchmark derangement for irrelevant sources, a breakage guard of 100 Polish-correct items and shared entity spans. P0 (02, no hooks) scores the three models with the scaffold language swapped and with the English twin's prompt prepended to the Polish prompt. 03 builds the patch: the English twin's residual-stream state entering a source decoder layer at an anchor replaces, without scaling, the Polish twin's state entering a target layer while its options are scored; its smoke gates are a bitwise self-identity check, a byte-identical 10-item repeat and a calibration of every benchmark and anchor. P1 (04) scans the subsample at the scaffold-final anchor over every second layer as source and target, plus the pairs 23 into 13 and 3 into 13. P2 confirms the top P1 pairs (at most 6, all within 2 percentage points of the best) on all forward-discordant pairs with an irrelevant-source control, the breakage guard, and the anchor ladder: scaffold-final token, question-final token, shared entity span (mean-pooled), a random position and the first token. P3 runs the reverse direction, the two continued-pretraining arms at the P2 pairs, and a STEM slice of P2. 05 computes item bootstrap intervals (2000 resamples, seed 20260703) and exact McNemar tests.
Dataset
Translation pairs of 600 Belebele test items (English and Polish) and 600 MMLU test items with Polish machine translations from openGPT-X mmlux, with committed per-item scores of Qwen2.5-1.5B and its two continued-pretraining arms; the patch sets exclude structurally malformed pairs.
Split
None: confirmation patches every forward-discordant pair; the coarse scan uses a seeded subsample of 100 (50 per benchmark).
Access needs
The two continued-pretraining checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer reference used for the APT4 arm is gated on Hugging Face.
Configurations
Anchors: scaffold-final token (primary), question-final token, shared entity span, random position, first token. Layers: every second decoder layer as source and target, plus 23 into 13 and 3 into 13. Sources: the item's own English twin (specific) and another item's (irrelevant). Models: Qwen2.5-1.5B; continued-pretraining arms A and B at 500M tokens. Behavioural conditions: baseline, scaffold swap, translation assist.
Metric
Flip rate: share of forward-discordant pairs whose patched Polish prediction is correct; irrelevant-source flip rate; breakage on Polish-correct pairs; translation-assist recovery of the English-minus-Polish accuracy gap; accuracy change under a scaffold-language swap.
Seeds
20260703 (PCG64: scan subsample, derangement, breakage guard, bootstrap)
Interpretation rule
Validity gate first: specific flip rate at least 3 times the irrelevant-source flip rate, and breakage at most 5 percentage points on the Polish-correct guard; if it fails, patching claims are void and the experiment is reported as an instrument failure. Then the flip rate R at the best confirmed pair: R at least 0.50 is routing-dominant, R at most 0.15 is representation-bound, which also requires the entity-span anchor to fail, otherwise mixed, reported by benchmark. Translation-assist recovery of the gap on the base model, per benchmark: at least 70% comprehension-dominant, at most 30% deeper than comprehension, otherwise mixed. Scaffold swap: absolute accuracy change below 3 percentage points means the gap tracks question language. Arm A's flip rate within 10 percentage points of the base model's with an interval including zero means continued pretraining left routing unchanged. Reverse flip rate at most half the forward rate means an English-sided store. Fixed pairs 23 into 13 and 3 into 13, read only if routing is dominant or mixed: at least 50% of the best R. STEM against non-STEM flip rate is exploratory, without a bin.
Resources
Apple silicon, 128 GB unified memory; bf16 inference with forward pre-hooks on the MPS backend, about 5 to 7 hours; no training.
Prior work
arXiv:2607.08393 patches residual states between layers of the same failing prompt (self-patching), reports effective pairs from about 0.10L and 0.82L into 0.45L, and uses an irrelevant-patching control.

Selected exact hypotheses and premises

Hypothesis · H34

cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Whether the gap tracks question language rather than scaffold format.

Hypothesis · H35

489055f5-af11-44e3-b63d-c59d6d7cfef6

How much of the gap the English twin in the prompt recovers.

Hypothesis · H36

cd49225b-c235-425c-83bc-e24271555357

The keystone prediction: routing-bound failures reachable by the patch.

Hypothesis · H37

fe0d162e-9992-4466-9b39-c6fb4a6ff108

Where effective layer pairs lie.

Hypothesis · H38

72ecb52b-0fef-4a96-8a5b-4ec81fb1e202

Which token positions carry the effect.

Hypothesis · H39

e73869df-b1f8-4941-873f-5b3ba6931164

Whether continued pretraining changes the patch effect.

Hypothesis · H40

764aaab5-717b-4376-abb9-30c3788c94b1

Whether the effect is asymmetric between directions.

Hypothesis · H41

69a5ef7f-14e9-4856-a165-fffa8dec769e

Whether two fixed layer pairs suffice.

Premise · P244

e6ce775f-1927-4ab2-90fc-cdab7bcec624

The English-minus-Polish accuracy gap of the base model on these translation pairs.

Premise · P247

22e9d053-dd03-4cb2-ba2d-bea9e71a1260

The English-minus-Polish accuracy gap of the base model on these translation pairs.

Premise · P264

72f15189-34bd-468a-a7dd-44743ac5531b

Continued pretraining bought no Polish access gain on the same models.