Execution attempt

Sign in with GitHub
← Experiment E12 · When Qwen2.5-1.5B answers a translation-paired multiple-choice item correctly in English and incorrectly in Polish, is the knowledge present in its forward pass but not reached from the Polish prompt, or absent from what the Polish prompt can reach?

Execution attempt · A120 · planned

Re-scores the English and Polish baseline cells of Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer and checks them against the committed per-item scores.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 7253735a8d117a39ade0ca39443d0726d77919fc

Reference checked 2026-09-14 13:18 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/02_behavioral.py --tag armA_500M --condition baseline
Working directory
experiments/E12-activation-patching
Configuration paths
None
Parameters
model armA_500M; condition baseline; 600 pairs per benchmark; comparison at 4 decimals and predicted letter
Environment
3.12.13; PyTorch MPS, bf16, decoder-layer forward pre-hooks; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • original-tokenizer arm at 500M tokens · model · M140 ckpt_00500M · raw output · Ask the reporter
  • E11 belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · raw output · Download
  • E11 belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · raw output · Download
  • E11 mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · raw output · Download
  • E11 mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · raw output · Download
  • E11 scores_armA_500M.jsonl · reference scores · M502 scores_armA_500M.jsonl as run · raw output · Rebuild via attempt

Delivered events

  1. Registered

    #1

    Re-scores the English and Polish baseline cells of Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer and checks them against the committed per-item scores.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • behavioral_baseline_armA_500M.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/633b3ba1ce916e432ae5ba4e0a706ad58b53b2e9/experiments/E12-activation-patching/results/behavioral_baseline_armA_500M.jsonl · public · sha256 5d4b673982ae… · 483970 bytes · M520
    • identity gate output armA_500M · Not kept: the gate's printed result was not committed · unavailable

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.