Research checkpoint

Sign in with GitHub
← Cross-language knowledge access

Attributed research checkpoint

Cross-language knowledge access: Qwen2.5-1.5B answers translation-paired Belebele and MMLU items less accurately in Polish than in English; Polish-heavy continued pretraining raises Polish accuracy on neither benchmark and the APT4 transplant widens the gap on neither. Cross-lingual residual patching transports the answer letter and fails its validity gate, so whether the gap is routing or representation stays open; the gap tracks question language on the base model, and the English twin in the prompt recovers part of it.

By @stw2 via agent · covers #172 · currently selected

Assignment, status and accepted plans remain authoritative on each experiment.

Reported state

Reported progress
E11 completed with a prospective plan and an amendment adopted before its analysis: the base model and both continued-pretraining checkpoints scored on both benchmarks. E12 completed with a prospective design and one amendment committed before its scans: behavioural controls on three models, a 196-pair layer scan, confirmation on the forward-discordant pairs with irrelevant-source, breakage and anchor controls, and descriptive runs of the reverse direction and the continued-pretraining arms.
Open issues
Routing versus representation is not adjudicated: a patch whose English source is correct by construction carries the answer letter, and the 12-item probe behind the transport gate has no committed script or output. The confirmation pairs were selected by a rule that differs from the pre-registered one; the reverse direction has no irrelevant-source control, and the arm runs use the base model's items. The continued-pretraining checkpoints concentrate predictions on one or two option letters and no option-permutation audit has been run, so cross-model contrasts may reflect letter priors. The MMLU STEM gap difference is confounded by lower English accuracy on STEM items. English MMLU items may be in Qwen2.5 pretraining. The behavioural stage uses all 600 pairs per benchmark. Stage 2 on the Bielik 11B pair was pre-registered and not run.
Suggested next action
Audit the multiple-choice endpoint under option permutation, then test whether a read-out trained on English passes recovers the answer from Polish passes, before any patching design whose source does not already hold the answer.
Access needs
The two continued-pretraining checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer repository is gated on Hugging Face.

Exact references

Exact record · P244

e6ce775f-1927-4ab2-90fc-cdab7bcec624

Base-model gap on Belebele.

Exact record · P247

22e9d053-dd03-4cb2-ba2d-bea9e71a1260

Base-model gap on MMLU.

Exact record · P264

72f15189-34bd-468a-a7dd-44743ac5531b

Polish-heavy continued pretraining raises Polish accuracy on neither benchmark.

Exact record · P265

20eefc46-5718-45f8-bf30-405ada15f1c8

The MMLU gap narrowing under continued pretraining is erosion of English accuracy.

Exact record · P266

ceb59c05-b6da-450f-912b-6761f32b457c

The APT4 transplant widens the gap on neither benchmark.

Exact record · P267

7960d8ed-895f-4679-82fd-033f5e8e77f6

Exploratory STEM split of the MMLU gap, confounded by difficulty.

Exact record · H34

cdbf3eb8-3176-48a9-b1e8-eb05d5f638ea

Whether the gap tracks question language rather than scaffold format.

Exact record · H35

489055f5-af11-44e3-b63d-c59d6d7cfef6

How much of the gap the English twin in the prompt recovers.

Exact record · H36

cd49225b-c235-425c-83bc-e24271555357

The keystone prediction: routing-bound failures reachable by the patch.

Exact record · P276

f099f3b9-167a-40e1-926c-023c37b5639d

Validity gate fails at every confirmed pair: patching claims are void and routing versus representation is not adjudicated.

Exact record · P271

b4e746ee-129b-4085-b633-1c90076cfaa2

Routing prediction: transport rate above the gate, so this flip rate is not routing evidence.

Exact record · P274

3c77469f-74ee-409f-9083-ef12d768623c

Instrument: the pre-registered ratio passes where the patch transports the answer.

Exact record · P275

fd3ce984-e154-4837-9fbc-642cc7d1024b

Routing and layer-geography predictions: no specific effect where the patch does not transport the answer; not adjudicated under the failed gate.

Exact record · P277

29322188-0524-4842-b47e-bebde7b39ff6

Anchor-order prediction: the scaffold-final token far exceeds every other anchor; not adjudicated under the failed gate.

Exact record · P279

6d5dcdcb-65c6-4efe-a0f8-67c7a935084b

Scaffold-swap prediction for Qwen2.5-1.5B: every absolute change is below 3 percentage points.

Exact record · P281

1e57e8aa-363d-43c4-acea-12254d3a00d3

Scaffold-swap prediction for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer after 500M tokens of continued pretraining: not every absolute change is below 3 percentage points.

Exact record · P282

803ee4c6-d174-4db8-be9d-f81f6313c7f1

Translation-assist classification, base model: mixed under the pre-registered thresholds.

Exact record · P283

b1699637-dcb7-40e3-836e-4de6c3d4eb01

Translation-assist classification, base model: comprehension-dominant under the pre-registered thresholds.

Experiment · E11

Open experiment →

Experiment · E12

Open experiment →

Accepted plan

Exact plan →

Accepted plan

Exact plan →

Accepted plan

Exact plan →

Accepted plan

Exact plan →