Execution attempt

Sign in with GitHub
← Experiment E4 · At equal character budgets of scientific documents, does long-context task accuracy fall earlier for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct in English and later in Polish, and is the crossover where the token-window arithmetic puts it?

Execution attempt · A32 · planned

Counts generations containing the reasoning tag per model and language.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ c60abf85776b3d991bdbba0e47569c6d944056e7

Reference checked 2026-09-14 13:10 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/05_think_rates.py
Working directory
experiments/E04-effective-context
Configuration paths
None
Parameters
literal tag <think>; grid generations only
Environment
3.14.5; CPU, standard library only; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • run_pl.jsonl · raw output · M115 run_pl.jsonl · raw output · Ask the reporter
  • run_orig.jsonl · raw output · M116 run_orig.jsonl · raw output · Ask the reporter

Delivered events

  1. Registered

    #1

    Counts generations containing the reasoning tag per model and language.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • think_rates.json · https://github.com/stw2/tokenizer-science-tax/blob/c60abf85776b3d991bdbba0e47569c6d944056e7/experiments/E04-effective-context/results/think_rates.json · public · sha256 e3eab8d8744a… · 349 bytes · M122

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.