Execution attempt

Sign in with GitHub
← Experiment E3 · Does forcing English chain-of-thought on Polish STEM questions raise the accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct?

Execution attempt · A24 · planned

Writes flat headline values from the analysis, the token-cap audit and the item manifest.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ d3c436d89b85172cdba28025d7df1f31126c8db2

Reference checked 2026-09-14 13:10 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/07_metrics.py
Working directory
experiments/E03-reasoning-language
Configuration paths
None
Parameters
none
Environment
3.14.5; CPU, standard library; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

dd_all
-0.0225
dd_pes
-0.036
items_pes
250
dd_gsm8k_pl
-0.036
items_total
800
dd_pes_holm_p
0.778
items_complete
800
items_gsm8k_pl
250
dd_all_ci95_low
-0.055
dd_llmzszl_stem
0
dd_pes_ci95_low
-0.116
dd_all_ci95_high
0.0112
dd_pes_ci95_high
0.044
dd_all_bootstrap_p
0.191
dd_gsm8k_pl_holm_p
0.429
dd_pes_bootstrap_p
0.389
items_llmzszl_stem
300
pes_specializations
57
dd_gsm8k_pl_ci95_low
-0.08
dd_gsm8k_pl_ci95_high
0.012
gsm8k_pl_replacements
0
dd_llmzszl_stem_holm_p
1
dd_truncation_excluded
-0.026
en_cot_gain_bielik_11b
0.0088

and 97 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    Writes flat headline values from the analysis, the token-cap audit and the item manifest.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    dd_all=-0.0225 · dd_pes=-0.036 · items_pes=250 · dd_gsm8k_pl=-0.036 · items_total=800 · dd_pes_holm_p=0.778 · items_complete=800 · items_gsm8k_pl=250 · …

    • metrics.json · https://github.com/stw2/tokenizer-science-tax/blob/0442df47c4492aa995ece02fa9b8dbc5e6df365f/experiments/E03-reasoning-language/results/metrics.json · public · sha256 10b442a23cba… · values 394858ea9aec… · 5474 bytes · M108

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.