Experiment

Sign in with GitHub
← Experiments

Experiment · E18

When Qwen2.5-1.5B answers a translation pair correctly in English and incorrectly in Polish, can a linear read-out trained on English passes recover the gold option from the Polish pass?

Available · Proposed by @stw2 · Unassigned

Qwen2.5-1.5B (primary) and its two continued-pretraining arms after 500M tokens, on MMLU translation pairs (primary) and Belebele translation pairs, with forward-discordant, both-correct and both-wrong strata from the option-permutation audit's permutation-marginalised scores. Linear read-outs (ridge or logistic) that score whether an option is the gold option from residual-stream states at each option's final token: trained on English or Polish passes, tested on English or Polish passes, at every decoder layer, with train and test split over items. Inference and linear probes only; nothing is inserted into the forward pass.

Prerequisites and protocol

Access needs
The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. Runs on Apple silicon, 128 GB unified memory.
Suggested protocol
Precondition: the option-permutation audit. Commit a design before extracting states. Prediction per item: the option with the highest read-out score. The both-wrong stratum holds the representation source fixed and lacks the knowledge, so decodability on both-wrong pairs equal to that on discordant pairs shows the read-out solving the item from the input; a layer-0 embedding baseline bounds lexical overlap; a control task with randomly assigned targets measures selectivity. Answer transport is impossible because nothing is inserted. Kill criterion: transfer accuracy on forward-discordant pairs at most that on both-wrong pairs, which reads the failure as representation-bound. Residual-stream hooks follow scripts/03_patch_harness.py and scripts/e12lib.py of experiments/E12-activation-patching.

Accepted plan

No accepted plan. Available work need not have a complete protocol or source commit.

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

No attempt registered. Work status and findings are independent of attempts.

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →