Execution attempt · A119 · planned
Re-scores the English and Polish baseline cells of Qwen2.5-1.5B and checks them against the committed per-item scores.
Pinned source and configuration
https://github.com/stw2/tokenizer-science-tax @ 7253735a8d117a39ade0ca39443d0726d77919fc
Reference checked 2026-09-14 13:18 UTC. Later commits, branches or plan changes do not retarget this attempt.
- Command
- python scripts/02_behavioral.py --tag base --condition baseline
- Working directory
- experiments/E12-activation-patching
- Configuration paths
- None
- Parameters
- model base; condition baseline; 600 pairs per benchmark; comparison at 4 decimals and predicted letter
- Environment
- 3.12.13; PyTorch MPS, bf16, decoder-layer forward pre-hooks; Apple silicon, 128 GB unified memory
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- Qwen2.5-1.5B · model · M124 Qwen2.5-1.5B · checkpoint · Download
- E11 belebele_en.jsonl · evaluation data · M485 belebele_en.jsonl · raw output · Download
- E11 belebele_pl.jsonl · evaluation data · M486 belebele_pl.jsonl · raw output · Download
- E11 mmlu_en.jsonl · evaluation data · M487 mmlu_en.jsonl · raw output · Download
- E11 mmlu_pl.jsonl · evaluation data · M488 mmlu_pl.jsonl · raw output · Download
- E11 scores_base.jsonl · reference scores · M496 scores_base.jsonl · raw output · Download
Delivered events
Registered
#1Re-scores the English and Polish baseline cells of Qwen2.5-1.5B and checks them against the committed per-item scores.
Started
#2Started.
Succeeded
#3Completed with exit code 0.
Exit code 0.
- behavioral_baseline_base.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/633b3ba1ce916e432ae5ba4e0a706ad58b53b2e9/experiments/E12-activation-patching/results/behavioral_baseline_base.jsonl · public · sha256 d9ff0d896053… · 471691 bytes · M519
- identity gate output base · Not kept: the gate's printed result was not committed · unavailable
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.