Execution attempt

Sign in with GitHub
← Experiment E6 · Does a math and code share in recovery pretraining of an APT4 transplant remove its regression on arithmetic text and digit arithmetic at a fixed token budget, and what does it cost on Polish text?

Execution attempt · A56 · planned

Scores Qwen2.5-0.5B on the ten evaluation texts with PyTorch on Apple silicon, for comparison with the CUDA rows (gate G1).

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 1e3dd3f736856c404f72ee42eb26b08b2062c223

Reference checked 2026-09-14 13:12 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/05_score_bpb.py --model Qwen/Qwen2.5-0.5B --splits pl=../E05-control-arm/data/raw/pl_eval.jsonl en=../E05-control-arm/data/raw/en_eval.jsonl sci_arxiv=../E01-fertility-atlas/data/corpora/en-arxiv-abstracts.jsonl sci_latex=../E01-fertility-atlas/data/corpora/en-latex-methods.jsonl sci_python=../E01-fertility-atlas/data/corpora/en-python-code.jsonl sci_gsm8k=../E01-fertility-atlas/data/corpora/en-gsm8k.jsonl pl_wiki_sci=../E01-fertility-atlas/data/corpora/pl-wiki-science.jsonl pl_pes=../E01-fertility-atlas/data/corpora/pl-science-pes.jsonl pl_informal=../E01-fertility-atlas/data/corpora/pl-informal.jsonl math_clean=../E05-control-arm/data/eval/math_clean.jsonl --max-docs 200 --out results/bpb_g1_mac.jsonl
Working directory
experiments/E06-math-code-rescue
Configuration paths
None
Parameters
first 200 documents per text; windows of 2,048 tokens; base model only
Environment
not recorded; PyTorch MPS, bf16; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • Qwen2.5-0.5B · base model · M189 Qwen2.5-0.5B base model · model · Download
  • Polish FineWeb2-HQ evaluation holdout · evaluation data · M129 pl_eval.jsonl · raw output · Download
  • English SlimPajama-6B evaluation holdout · evaluation data · M130 en_eval.jsonl · raw output · Ask the reporter
  • math_clean synthetic arithmetic set · evaluation data · M128 math_clean evaluation statements · dataset · Download
  • en-arxiv-abstracts.jsonl · evaluation data · M27 en-arxiv-abstracts.jsonl · raw output · Download
  • en-latex-methods.jsonl · evaluation data · M28 en-latex-methods.jsonl · raw output · Ask the reporter
  • en-python-code.jsonl · evaluation data · M29 en-python-code.jsonl · raw output · Ask the reporter
  • en-gsm8k.jsonl · evaluation data · M30 en-gsm8k.jsonl · raw output · Download
  • pl-science-pes.jsonl · evaluation data · M31 pl-science-pes.jsonl · raw output · Ask the reporter
  • pl-wiki-science.jsonl · evaluation data · M32 pl-wiki-science.jsonl · raw output · Download
  • pl-informal.jsonl · evaluation data · M33 pl-informal.jsonl · raw output · Ask the reporter

Delivered events

  1. Registered

    #1

    Scores Qwen2.5-0.5B on the ten evaluation texts with PyTorch on Apple silicon, for comparison with the CUDA rows (gate G1).

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • bpb_g1_mac.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/bpb_g1_mac.jsonl · public · sha256 4bc814f63582… · 1181 bytes · M283

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.