Experiment · E11
Does Qwen2.5-1.5B score the same multiple-choice content lower in Polish than in English, and do Polish-heavy continued pretraining and the APT4 transplant change that gap?
Qwen2.5-1.5B and its two checkpoints after 500M tokens of Polish-heavy continued pretraining, one keeping the original tokenizer and one with APT4 by FVT, on 600 translation-paired Belebele and 600 translation-paired MMLU items; likelihood scoring only, no training.
Prerequisites and protocol
- Access needs
- The two continued-pretraining checkpoints are restricted materials; the APT4 tokenizer repository is gated.
- Suggested protocol
- Scripts 01 to 05 in experiments/E11-cross-language-access, in order.
Accepted plan
Plan accepted by @stw2
Paired English and Polish likelihood multiple-choice scoring of Qwen2.5-1.5B and two continued-pretraining checkpoints on Belebele and MMLU- requires Belebele eng_Latn test · item source · M481 Belebele English test split · Download
- requires Belebele pol_Latn test · item source · M482 Belebele Polish test split · Download
- requires MMLU all test · item source · M483 MMLU all test split · Download
- requires openGPT-X mmlux Polish test files · item source · M484 openGPT-X MMLU translations · Download
- requires Qwen2.5-1.5B · checkpoint · M124 Qwen2.5-1.5B · Download
- requires checkpoint armA ckpt_00500M · checkpoint · M140 ckpt_00500M · Ask the reporter
- requires checkpoint armB ckpt_00500M · checkpoint · M145 ckpt_00500M · Ask the reporter
- requires APT4 reference tokenizer.json · tokenizer · M125 APT4 tokenizer.json (Bielik-PL-11B-v3.0-Instruct) · Ask the reporter
- requires belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · Download
- requires belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · Download
- requires mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · Download
- requires mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · Download
- requires exclusions.json · configuration · M515 exclusions.json (13 excluded item pairs) · Download
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 9
9 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
A109 Builds the translation-paired Belebele and MMLU item files from the seeded source samples and gates pairing and gold parity.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan
A114 Scores all 2,400 English and Polish items with the continued-pretraining checkpoint that kept the original tokenizer.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan
A115 Scores all 2,400 English and Polish items with the continued-pretraining checkpoint that has APT4 by FVT.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 3c128813c117 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 5e8117ffb3f1 · under an accepted plan
A117 Writes flat headline values from the analysis and counts letters and subset accuracies from the score files.
Succeededplanned attempt · stw2/tokenizer-science-tax @ a544bf7f4ee7 · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Cross-language knowledge access. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →