Execution attempt

Sign in with GitHub
← Experiment E12 · When Qwen2.5-1.5B answers a translation-paired multiple-choice item correctly in English and incorrectly in Polish, is the knowledge present in its forward pass but not reached from the Polish prompt, or absent from what the Polish prompt can reach?

Execution attempt · A124 · planned

Scores Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer on both translation-pair sets with the answer scaffold in the other language.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 7253735a8d117a39ade0ca39443d0726d77919fc

Reference checked 2026-09-14 13:18 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/02_behavioral.py --tag armA_500M --condition xscaffold
Working directory
experiments/E12-activation-patching
Configuration paths
None
Parameters
model armA_500M; condition xscaffold; 600 pairs per benchmark; no hooks
Environment
3.12.13; PyTorch MPS, bf16, decoder-layer forward pre-hooks; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • original-tokenizer arm at 500M tokens · model · M140 ckpt_00500M · raw output · Ask the reporter
  • E11 belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · raw output · Download
  • E11 belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · raw output · Download
  • E11 mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · raw output · Download
  • E11 mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · raw output · Download

Delivered events

  1. Registered

    #1

    Scores Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer on both translation-pair sets with the answer scaffold in the other language.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • behavioral_xscaffold_armA_500M.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/633b3ba1ce916e432ae5ba4e0a706ad58b53b2e9/experiments/E12-activation-patching/results/behavioral_xscaffold_armA_500M.jsonl · public · sha256 ca4cd98a9180… · 505549 bytes · M524

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.