Execution attempt

Sign in with GitHub
← Experiment E11 · Does Qwen2.5-1.5B score the same multiple-choice content lower in Polish than in English, and do Polish-heavy continued pretraining and the APT4 transplant change that gap?

Execution attempt · A117 · planned

Writes flat headline values from the analysis and counts letters and subset accuracies from the score files.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ a544bf7f4ee73f31ddaa43cf5b1218457461d94e

Reference checked 2026-09-14 13:18 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/04_metrics.py
Working directory
experiments/E11-cross-language-access
Configuration paths
None
Parameters
none
Environment
3.14.5; Python standard library; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

cells_total
12
mmlu_dg_cpt
-0.10356536502546693
mmlu_n_pairs
589
mmlu_base_gap
0.21392190152801357
mmlu_dg_trans
-0.01528013582342952
h4_mmlu_n_stem
111
belebele_dg_cpt
-0.020066889632106955
belebele_n_pairs
598
mmlu_base_en_acc
0.5568760611205433
mmlu_base_pl_acc
0.34295415959252973
mmlu_dacc_en_cpt
-0.10356536502546693
mmlu_dacc_pl_cpt
0
mmlu_floor_cells
0
belebele_base_gap
0.19732441471571907
belebele_dg_trans
-0.04180602006688966
gold_count_mmlu_a
134
gold_count_mmlu_b
159
gold_count_mmlu_c
147
gold_count_mmlu_d
160
h4_mmlu_n_nonstem
478
mmlu_dacc_en_trans
0.00509337860780984
mmlu_dacc_pl_trans
0.02037351443123936
rows_per_file_mmlu
600
mmlu_arm_a_500m_gap
0.11035653650254669

and 161 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    Writes flat headline values from the analysis and counts letters and subset accuracies from the score files.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    cells_total=12 · mmlu_dg_cpt=-0.10356536502546693 · mmlu_n_pairs=589 · mmlu_base_gap=0.21392190152801357 · mmlu_dg_trans=-0.01528013582342952 · h4_mmlu_n_stem=111 · belebele_dg_cpt=-0.020066889632106955 · belebele_n_pairs=598 · …

    • metrics.json · https://github.com/stw2/tokenizer-science-tax/blob/5e8117ffb3f168331ee63f1032cd40a62c6296bd/experiments/E11-cross-language-access/results/metrics.json · public · sha256 c6b1c7d198db… · values 22ef974eeffb… · 8791 bytes · M517

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.