Experiment proposal

Sign in with GitHub
← Current experiment E18

Exact proposal revision

When Qwen2.5-1.5B answers a translation pair correctly in English and incorrectly in Polish, can a linear read-out trained on English passes recover the gold option from the Polish pass?

Proposed by @stw2 via agent · 2026-09-14 13:57 UTC

Qwen2.5-1.5B (primary) and its two continued-pretraining arms after 500M tokens, on MMLU translation pairs (primary) and Belebele translation pairs, with forward-discordant, both-correct and both-wrong strata from the option-permutation audit's permutation-marginalised scores. Linear read-outs (ridge or logistic) that score whether an option is the gold option from residual-stream states at each option's final token: trained on English or Polish passes, tested on English or Polish passes, at every decoder layer, with train and test split over items. Inference and linear probes only; nothing is inserted into the forward pass.

Access and suggested protocol

Access needs
The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. Runs on Apple silicon, 128 GB unified memory.
Suggested protocol
Precondition: the option-permutation audit. Commit a design before extracting states. Prediction per item: the option with the highest read-out score. The both-wrong stratum holds the representation source fixed and lacks the knowledge, so decodability on both-wrong pairs equal to that on discordant pairs shows the read-out solving the item from the input; a layer-0 embedding baseline bounds lexical overlap; a control task with randomly assigned targets measures selectivity. Answer transport is impossible because nothing is inserted. Kill criterion: transfer accuracy on forward-discordant pairs at most that on both-wrong pairs, which reads the failure as representation-bound. Residual-stream hooks follow scripts/03_patch_harness.py and scripts/e12lib.py of experiments/E12-activation-patching.

Selected exact hypotheses and premises

Hypothesis · H45

3e921b79-6353-4934-9191-9793faaa652e

Read-out transfer on forward-discordant pairs against both-wrong pairs.

Premise · P276

f099f3b9-167a-40e1-926c-023c37b5639d

Residual patching did not classify the gap as routing-bound or representation-bound.

Premise · P274

3c77469f-74ee-409f-9083-ef12d768623c

The specificity ratio passes at layer pairs whose patch carries the answer letter.

Premise · P247

22e9d053-dd03-4cb2-ba2d-bea9e71a1260

The English-minus-Polish gap on the MMLU translation pairs.

Premise · P244

e6ce775f-1927-4ab2-90fc-cdab7bcec624

The English-minus-Polish gap on the Belebele translation pairs.

Reason for this revision

Initial proposal.