Execution attempt

Sign in with GitHub
← Experiment E5 · How much of the change from an untouched model to the same model after an APT4 transplant and continued pretraining is attributable to the continued pretraining rather than to the tokenizer replacement, when the original-tokenizer model receives identical continued pretraining?

Execution attempt · A34 · planned

Streams FineWeb2-HQ Polish and SlimPajama-6B English with a fixed shuffle, holds out 2,000 documents of each and counts tokens with the Qwen and APT4 tokenizers.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 4d64dedcbe57761c48ea37d5ce80cb6c86198236

Reference checked 2026-09-14 13:10 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/01_download_data.py
Working directory
experiments/E05-control-arm
Configuration paths
None
Parameters
Polish target 960e6 and English target 240e6 Qwen tokens; 2,000 held-out documents per source; streaming shuffle seed 42 with a 10,000-document buffer; documents under 200 characters skipped. The run loaded the default branches; the pins are inferred.
Environment
unrecorded; CPU tokenization, datasets streaming; one CUDA GPU, 16 GB
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Delivered events

  1. Registered

    #1

    Streams FineWeb2-HQ Polish and SlimPajama-6B English with a fixed shuffle, holds out 2,000 documents of each and counts tokens with the Qwen and APT4 tokenizers.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • pl_eval.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/4d64dedcbe57761c48ea37d5ce80cb6c86198236/experiments/E05-control-arm/data/raw/pl_eval.jsonl · public · sha256 6d833121c4be… · 5932759 bytes · M129
    • en_eval.jsonl · Not redistributed: ask the Room owner, or rebuild with scripts/01_download_data.py at the pinned revisions · restricted · sha256 6a16eae06eb1… · 8068309 bytes · M130
    • manifest.json · https://github.com/stw2/tokenizer-science-tax/blob/4d64dedcbe57761c48ea37d5ce80cb6c86198236/experiments/E05-control-arm/data/raw/manifest.json · public · sha256 12aff25f1b1e… · 639 bytes · M131
    • pl.jsonl · Not kept: written on the original training machine and never copied; rebuild with scripts/01_download_data.py at the pinned revisions · unavailable
    • en.jsonl · Not kept: written on the original training machine and never copied; rebuild with scripts/01_download_data.py at the pinned revisions · unavailable

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.