Experiment · E19
When the Polish pass's options are permuted independently of its English twin's, does patching the English twin's residual state make the Polish pass choose its own gold option, or the letter of the English twin's gold?
Qwen2.5-1.5B and, as secondary, its two continued-pretraining arms, each on its own forward-discordant translation pairs from the option-permutation audit. Permuted cross-lingual patch: the Polish option block re-rendered in an independently drawn order, reporting S, the rate of predicting the English twin's gold letter, and C, the rate of predicting the Polish pass's own gold option, on pairs whose two gold letters differ, with coinciding-letter pairs reported separately and never pooled. Self-patching: source and target within the same failing Polish pass, from a source layer to a target layer. Option-free source: an English source with passage and question only. The breakage guard is re-run under permutation and the reverse direction gets an irrelevant-source control. Inference only.
Prerequisites and protocol
- Access needs
- The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. Runs on Apple silicon, 128 GB unified memory.
- Suggested protocol
- Preconditions: the option-permutation audit, and the English-trained read-out test run first; this experiment is the causal check if the read-out finds the answer present but unused. Commit a design, with the mutant walk, before scanning. Under permutation S and C name different letters, so a patch that carries only an answer letter can raise only S, and C cannot be raised by lowering an irrelevant-source floor. Self-patching inserts nothing correct, so a flip there is internal re-routing. New code: an option parser for the item files of experiments/E11-cross-language-access and a permuted re-render of each prompt; the patch harness of experiments/E12-activation-patching is otherwise reused. Kill criteria: fewer than 60% of pairs stable under permutation, or C on coinciding-letter pairs equal to C on the other pairs; either voids the permuted patch.
Accepted plan
No accepted plan. Available work need not have a complete protocol or source commit.
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
No attempt registered. Work status and findings are independent of attempts.
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →