Experiment

Sign in with GitHub
← Experiments

Experiment · E11

Does Qwen2.5-1.5B score the same multiple-choice content lower in Polish than in English, and do Polish-heavy continued pretraining and the APT4 transplant change that gap?

Completed · Proposed by @stw2 · Assigned to @stw2

Qwen2.5-1.5B and its two checkpoints after 500M tokens of Polish-heavy continued pretraining, one keeping the original tokenizer and one with APT4 by FVT, on 600 translation-paired Belebele and 600 translation-paired MMLU items; likelihood scoring only, no training.

Prerequisites and protocol

Access needs
The two continued-pretraining checkpoints are restricted materials; the APT4 tokenizer repository is gated.
Suggested protocol
Scripts 01 to 05 in experiments/E11-cross-language-access, in order.

Accepted plan

Plan accepted by @stw2

Paired English and Polish likelihood multiple-choice scoring of Qwen2.5-1.5B and two continued-pretraining checkpoints on Belebele and MMLU

retrospective

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts · 9

9 succeeded

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

  1. planned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  2. rerun attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  3. planned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  4. rerun attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  5. planned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  6. planned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  7. planned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  8. planned attempt · stw2/tokenizer-science-tax @ 5e8117ffb3f1 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  9. planned attempt · stw2/tokenizer-science-tax @ a544bf7f4ee7 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →