Execution attempt

Sign in with GitHub
← Experiment E11 · Does Qwen2.5-1.5B score the same multiple-choice content lower in Polish than in English, and do Polish-heavy continued pretraining and the APT4 transplant change that gap?

Execution attempt · A111 · planned

Determinism gate, first run: scores the first 10 items of each item file with Qwen2.5-1.5B and writes one prompt dump per file.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 3c128813c117fd41b46d3a9949c2c8751134df27

Reference checked 2026-09-14 13:18 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/02_probe_paired_mcq.py --model Qwen/Qwen2.5-1.5B --tag smoke --limit 10 --out results/smoke_run1.jsonl --dump-prompts
Working directory
experiments/E11-cross-language-access
Configuration paths
None
Parameters
limit 10 items per file; tag smoke
Environment
3.12.13; PyTorch, MPS, bf16; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • Qwen2.5-1.5B · checkpoint · M124 Qwen2.5-1.5B · checkpoint · Download
  • belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · raw output · Download
  • belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · raw output · Download
  • mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · raw output · Download
  • mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · raw output · Download

Delivered events

  1. Registered

    #1

    Determinism gate, first run: scores the first 10 items of each item file with Qwen2.5-1.5B and writes one prompt dump per file.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • smoke_run1.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/smoke_run1.jsonl · public · sha256 7f3234fd173d… · 11572 bytes · M491
    • prompt_dump_smoke_belebele_en.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_smoke_belebele_en.json · public · sha256 c623f1752df3… · 997 bytes · M492
    • prompt_dump_smoke_belebele_pl.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_smoke_belebele_pl.json · public · sha256 4cec129b8108… · 1103 bytes · M493
    • prompt_dump_smoke_mmlu_en.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_smoke_mmlu_en.json · public · sha256 6c43ba615618… · 1101 bytes · M494
    • prompt_dump_smoke_mmlu_pl.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/prompt_dumps/prompt_dump_smoke_mmlu_pl.json · public · sha256 7ede88d5b030… · 1194 bytes · M495

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.