Experiment · E12
When Qwen2.5-1.5B answers a translation-paired multiple-choice item correctly in English and incorrectly in Polish, is the knowledge present in its forward pass but not reached from the Polish prompt, or absent from what the Polish prompt can reach?
Qwen2.5-1.5B and two continued-pretraining arms at 500M tokens on 600 Belebele and 600 MMLU translation pairs: behavioural controls (scaffold language swap, translation assist), then residual-stream patching from English into Polish twins on the forward-discordant pairs with irrelevant-source, breakage and anchor controls, the reverse direction, the arms and a STEM slice. Inference only.
Prerequisites and protocol
- Access needs
- The two continued-pretraining checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer reference used for the APT4 arm is gated on Hugging Face.
- Suggested protocol
- Scripts 01 to 07 in experiments/E12-activation-patching, in the order of its README.
Accepted plan
Plan accepted by @stw2
Cross-lingual residual-stream patching from English into Polish twins of translation-paired multiple-choice items on Qwen2.5-1.5B, with behavioural controls and two continued-pretraining arms- requires Qwen2.5-1.5B · model · M124 Qwen2.5-1.5B · Download
- requires original-tokenizer arm at 500M tokens · model · M140 ckpt_00500M · Ask the reporter
- requires APT4 FVT arm at 500M tokens · model · M145 ckpt_00500M · Ask the reporter
- requires APT4 tokenizer.json · tokenizer · M125 APT4 tokenizer.json (Bielik-PL-11B-v3.0-Instruct) · Ask the reporter
- requires E11 belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · Download
- requires E11 belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · Download
- requires E11 mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · Download
- requires E11 mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · Download
- requires E11 exclusions.json · evaluation data · M515 exclusions.json (13 excluded item pairs) · Download
- requires E11 scores_base.jsonl · reference scores · M496 scores_base.jsonl · Download
- requires E11 scores_armA_500M.jsonl · reference scores · M502 scores_armA_500M.jsonl as run · Rebuild via attempt
- requires E11 scores_armB_500M.jsonl · reference scores · M508 scores_armB_500M.jsonl as run · Rebuild via attempt
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 20
20 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
A119 Re-scores the English and Polish baseline cells of Qwen2.5-1.5B and checks them against the committed per-item scores.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
A123 Scores Qwen2.5-1.5B on both translation-pair sets with the answer scaffold in the other language.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
A126 Scores Qwen2.5-1.5B on both translation-pair sets with the English twin's prompt prepended to each Polish prompt.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
A134 Computes every pre-registered endpoint and the transport rates, with item bootstrap intervals and exact McNemar tests.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
A135 Reruns the set derivation on the committed inputs; the output is byte-identical to the committed sets.json.
Succeededrerun attempt · stw2/tokenizer-science-tax @ 7253735a8d11 · under an accepted plan
A136 Reruns the analysis on the committed rows; the output is byte-identical to the committed analysis.json.
Succeededrerun attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
exploratory attempt · stw2/tokenizer-science-tax @ 633b3ba1ce91 · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →