Execution attempt

Sign in with GitHub
← Experiment E3 · Does forcing English chain-of-thought on Polish STEM questions raise the accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct?

Execution attempt · A18 · planned

Generates answers to all 800 items under Polish and English chain-of-thought with both models, Bielik-PL-11B-v3.0-Instruct first.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 02380f96e12e849fae306c5e9bc44f12c7e7a866

Reference checked 2026-09-14 13:10 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/02_run.py
Working directory
experiments/E03-reasoning-language
Configuration paths
None
Parameters
max_tokens 1280 for GSM8K-PL and 768 for multiple-choice; greedy; both models
Environment
3.12.13; MLX (Metal) bf16, greedy decoding; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Delivered events

  1. Registered

    #1

    Generates answers to all 800 items under Polish and English chain-of-thought with both models, Bielik-PL-11B-v3.0-Instruct first.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0. Duration 43696 s.

    • raw/run_bielik-pl-11b-v3.0-instruct.jsonl before extension · Overwritten in place by scripts/02b_extend_truncated.py; digest never recorded · unavailable
    • raw/run_bielik-11b-v3.0-instruct.jsonl before extension · Overwritten in place by scripts/02b_extend_truncated.py; digest never recorded · unavailable
    • prompt_dump_bielik-pl-11b-v3.0-instruct.json · https://github.com/stw2/tokenizer-science-tax/blob/0442df47c4492aa995ece02fa9b8dbc5e6df365f/experiments/E03-reasoning-language/results/prompt_dump_bielik-pl-11b-v3.0-instruct.json · public · sha256 0924e965786b… · 1235 bytes · M97
    • prompt_dump_bielik-11b-v3.0-instruct.json · https://github.com/stw2/tokenizer-science-tax/blob/0442df47c4492aa995ece02fa9b8dbc5e6df365f/experiments/E03-reasoning-language/results/prompt_dump_bielik-11b-v3.0-instruct.json · public · sha256 bddd806fb853… · 1231 bytes · M98

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.