Execution attempt

Sign in with GitHub
← Experiment E8 · Does Bielik-PL-11B-v3.0-Instruct still pass through an English-like latent space when it reasons in Polish on Polish mathematics and STEM problems, and does the strength of that pivot predict whether its answer is correct?

Execution attempt · A82 · planned

Writes flat headline values from the analysis and the lexicon report.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ a9246a60690bf1627e4cc56095f64e3307508bb2

Reference checked 2026-09-14 13:12 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/04_metrics.py
Working directory
experiments/E08-latent-pivot
Configuration paths
None
Parameters
none beyond the input files
Environment
3.14.5; standard library; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

holm_b
0.42
holm_a1
0
holm_a2
0
bench_pes_a1
0.2023
bench_pes_a2
-1.0875
bench_pes_a1_n
223
n_items_universe
766
b_coef_transplant
0.0907
bench_gsm8k_pl_a1
0.1709
bench_gsm8k_pl_a2
-0.5882
a2_pivot_norm_diff
-0.771
b_coef_transplant_n
766
bench_gsm8k_pl_a1_n
248
bench_pes_a2_p_holm
0
a2_pivot_excess_diff
-0.5098
a2_pivot_norm_diff_n
766
sens_a1_pooled_ratio
0.182
sens_b_tf_strict_098
0.1328
bench_llmzszl_stem_a1
0.171
bench_llmzszl_stem_a2
-0.6854
bench_pes_a1_ci95_low
0.1961
bench_pes_a2_ci95_low
-1.1008
lexicon_n_en_original
6116
lexicon_n_pl_original
1665

and 143 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    Writes flat headline values from the analysis and the lexicon report.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    holm_b=0.42 · holm_a1=0 · holm_a2=0 · bench_pes_a1=0.2023 · bench_pes_a2=-1.0875 · bench_pes_a1_n=223 · n_items_universe=766 · b_coef_transplant=0.0907 · …

    • metrics.json · https://github.com/stw2/tokenizer-science-tax/blob/a9246a60690bf1627e4cc56095f64e3307508bb2/experiments/E08-latent-pivot/results/metrics.json · public · sha256 7dc86b109fd3… · values c3fbf1a17a16… · 6876 bytes · M411

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.