Execution attempt

Sign in with GitHub
← Experiment E11 · Does Qwen2.5-1.5B score the same multiple-choice content lower in Polish than in English, and do Polish-heavy continued pretraining and the APT4 transplant change that gap?

Execution attempt · A113 · planned

Scores all 2,400 English and Polish items with Qwen2.5-1.5B.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 3c128813c117fd41b46d3a9949c2c8751134df27

Reference checked 2026-09-14 13:18 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/02_probe_paired_mcq.py --model Qwen/Qwen2.5-1.5B --tag base --out results/scores_base.jsonl --dump-prompts
Working directory
experiments/E11-cross-language-access
Configuration paths
None
Parameters
all items; tag base
Environment
3.12.13; PyTorch, MPS, bf16; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • Qwen2.5-1.5B · checkpoint · M124 Qwen2.5-1.5B · checkpoint · Download
  • belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · raw output · Download
  • belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · raw output · Download
  • mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · raw output · Download
  • mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · raw output · Download

Delivered events

  1. Registered

    #1

    Scores all 2,400 English and Polish items with Qwen2.5-1.5B.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • scores_base.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/scores_base.jsonl · public · sha256 96c21a7f99d4… · 690296 bytes · M496
    • prompt_dump_base_belebele_en.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_base_belebele_en.json · public · sha256 a676d217efe8… · 996 bytes · M497
    • prompt_dump_base_belebele_pl.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_base_belebele_pl.json · public · sha256 460932edf66b… · 1102 bytes · M498
    • prompt_dump_base_mmlu_en.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_base_mmlu_en.json · public · sha256 03960b35aab2… · 1100 bytes · M499
    • prompt_dump_base_mmlu_pl.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_base_mmlu_pl.json · public · sha256 25930c0e7dde… · 1193 bytes · M500

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.