Execution attempt · A127 · planned
Scores Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer on both translation-pair sets with the English twin's prompt prepended to each Polish prompt.
Pinned source and configuration
https://github.com/stw2/tokenizer-science-tax @ 7253735a8d117a39ade0ca39443d0726d77919fc
Reference checked 2026-09-14 13:18 UTC. Later commits, branches or plan changes do not retarget this attempt.
- Command
- python scripts/02_behavioral.py --tag armA_500M --condition assist
- Working directory
- experiments/E12-activation-patching
- Configuration paths
- None
- Parameters
- model armA_500M; condition assist; 600 pairs per benchmark; no hooks
- Environment
- 3.12.13; PyTorch MPS, bf16, decoder-layer forward pre-hooks; Apple silicon, 128 GB unified memory
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- original-tokenizer arm at 500M tokens · model · M140 ckpt_00500M · raw output · Ask the reporter
- E11 belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · raw output · Download
- E11 belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · raw output · Download
- E11 mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · raw output · Download
- E11 mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · raw output · Download
Delivered events
Registered
#1Scores Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer on both translation-pair sets with the English twin's prompt prepended to each Polish prompt.
Started
#2Started.
Succeeded
#3Completed with exit code 0.
Exit code 0.
- behavioral_assist_armA_500M.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/633b3ba1ce916e432ae5ba4e0a706ad58b53b2e9/experiments/E12-activation-patching/results/behavioral_assist_armA_500M.jsonl · public · sha256 05f463ff31ab… · 250349 bytes · M527
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.