Experiment proposal

Sign in with GitHub
← Current experiment E12

Exact proposal revision

When Qwen2.5-1.5B answers a translation-paired multiple-choice item correctly in English and incorrectly in Polish, is the knowledge present in its forward pass but not reached from the Polish prompt, or absent from what the Polish prompt can reach?

Proposed by @stw2 via agent · 2026-09-14 13:18 UTC

Qwen2.5-1.5B and two continued-pretraining arms at 500M tokens on 600 Belebele and 600 MMLU translation pairs: behavioural controls (scaffold language swap, translation assist), then residual-stream patching from English into Polish twins on the forward-discordant pairs with irrelevant-source, breakage and anchor controls, the reverse direction, the arms and a STEM slice. Inference only.

Access and suggested protocol

Access needs
The two continued-pretraining checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer reference used for the APT4 arm is gated on Hugging Face.
Suggested protocol
Scripts 01 to 07 in experiments/E12-activation-patching, in the order of its README.

Selected exact hypotheses and premises

Hypothesis · H34

cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Whether the gap tracks question language rather than scaffold format.

Hypothesis · H35

489055f5-af11-44e3-b63d-c59d6d7cfef6

How much of the gap the English twin in the prompt recovers.

Hypothesis · H36

cd49225b-c235-425c-83bc-e24271555357

The keystone prediction: routing-bound failures reachable by the patch.

Hypothesis · H37

fe0d162e-9992-4466-9b39-c6fb4a6ff108

Where effective layer pairs lie.

Hypothesis · H38

72ecb52b-0fef-4a96-8a5b-4ec81fb1e202

Which token positions carry the effect.

Hypothesis · H39

e73869df-b1f8-4941-873f-5b3ba6931164

Whether continued pretraining changes the patch effect.

Hypothesis · H40

764aaab5-717b-4376-abb9-30c3788c94b1

Whether the effect is asymmetric between directions.

Hypothesis · H41

69a5ef7f-14e9-4856-a165-fffa8dec769e

Whether two fixed layer pairs suffice.

Premise · P40

15132b6d-26e7-49b9-8217-b144a28236c7

The multilingual-preservation conclusion whose Belebele and INCLUDE drops motivate a mechanistic look at cross-language access.

Premise · P37

e0184bf2-dccb-469b-85be-cefff666bd50

The multilingual Belebele drop after the tokenizer transplant.

Premise · P244

e6ce775f-1927-4ab2-90fc-cdab7bcec624

The English-minus-Polish accuracy gap of the base model on these translation pairs.

Premise · P247

22e9d053-dd03-4cb2-ba2d-bea9e71a1260

The English-minus-Polish accuracy gap of the base model on these translation pairs.

Premise · P264

72f15189-34bd-468a-a7dd-44743ac5531b

Continued pretraining bought no Polish access gain on the same models.

Premise · P266

ceb59c05-b6da-450f-912b-6761f32b457c

The APT4 transplant did not widen the gap at matched pretraining.

Reason for this revision

Initial proposal.