Execution attempt

Sign in with GitHub
← Experiment E3 · Does forcing English chain-of-thought on Polish STEM questions raise the accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct?

Execution attempt · A19 · planned

Regenerates, for both models, every answer that stopped at its token cap with a 1280-token cap and rewrites the generation files.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ d3c436d89b85172cdba28025d7df1f31126c8db2

Reference checked 2026-09-14 13:10 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/02b_extend_truncated.py --cap 1280
Working directory
experiments/E03-reasoning-language
Configuration paths
None
Parameters
cap 1280; both models
Environment
3.12.13; MLX (Metal) bf16, greedy decoding; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Delivered events

  1. Registered

    #1

    Regenerates, for both models, every answer that stopped at its token cap with a 1280-token cap and rewrites the generation files.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0. Duration 6355 s.

    • raw/run_bielik-pl-11b-v3.0-instruct.jsonl · Not redistributed: ask the Room owner, or rebuild with scripts/02_run.py and scripts/02b_extend_truncated.py --cap 1280 at the pinned revisions · restricted · sha256 f1fb1688b747… · 1537126 bytes · M99
    • raw/run_bielik-11b-v3.0-instruct.jsonl · Not redistributed: ask the Room owner, or rebuild with scripts/02_run.py and scripts/02b_extend_truncated.py --cap 1280 at the pinned revisions · restricted · sha256 2badd92c308b… · 1239132 bytes · M100
    • extend_truncated.log · https://github.com/stw2/tokenizer-science-tax/blob/0442df47c4492aa995ece02fa9b8dbc5e6df365f/experiments/E03-reasoning-language/results/extend_truncated.log · public · sha256 9b2a3c00a48b… · 247 bytes · M101

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.