Execution attempt

Sign in with GitHub
← Experiment E4 · At equal character budgets of scientific documents, does long-context task accuracy fall earlier for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct in English and later in Polish, and is the crossover where the token-window arithmetic puts it?

Execution attempt · A33 · planned

Writes flat headline values from the build manifest, the graded records, the analysis and the reasoning-tag counts.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ c60abf85776b3d991bdbba0e47569c6d944056e7

Reference checked 2026-09-14 13:10 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/06_metrics.py
Working directory
experiments/E04-effective-context
Configuration paths
None
Parameters
none beyond the input files
Environment
3.14.5; CPU, standard library only; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • manifest.json · configuration · M112 manifest.json · raw output · Download
  • graded_pl.jsonl · analysis · M117 graded_pl.jsonl · raw output · Ask the reporter
  • graded_orig.jsonl · analysis · M118 graded_orig.jsonl · raw output · Ask the reporter
  • grade_summary.json · analysis · M119 grade_summary.json · raw output · Download
  • analysis.json · analysis · M120 analysis.json · raw output · Download
  • think_rates.json · analysis · M122 think_rates.json · raw output · Download

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

cells
690
controls
32
h1_ratio
1.4311097878431127
h3_chars
71237
h3_pairs
25
h1_p_holm
0.0014992503748125937
h2_p_holm
0.00228310502283105
h3_p_holm
1
h1_censored
1
h1_ci95_low
1.2850788483603277
h2_ci95_low
1.4997529469894826
h3_ci95_low
0
token_limit
32704
h1_ci95_high
1.588330165313058
h2_ci95_high
28.19400212314225
h3_ci95_high
0
h1_boot_valid
2000
h2_boot_valid
875
records_total
1217
h1_p_one_sided
0.0004997501249375312
h2_p_one_sided
0.001141552511415525
h3_p_one_sided
1
ceiling_ratio_en
1.4709
ceiling_ratio_pl
1.5324

and 153 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    Writes flat headline values from the build manifest, the graded records, the analysis and the reasoning-tag counts.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    cells=690 · controls=32 · h1_ratio=1.4311097878431127 · h3_chars=71237 · h3_pairs=25 · h1_p_holm=0.0014992503748125937 · h2_p_holm=0.00228310502283105 · h3_p_holm=1 · …

    • metrics.json · https://github.com/stw2/tokenizer-science-tax/blob/c60abf85776b3d991bdbba0e47569c6d944056e7/experiments/E04-effective-context/results/metrics.json · public · sha256 1d992ff49e6b… · values c68d2f8ba4bb… · 7813 bytes · M123

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.