Execution attempt · A83 · planned
Builds the LaTeX training text from the first proof-pile-2 arXiv train shard. Unlogged; the source flag is inferred from the manifest order and file times.
Pinned source and configuration
https://github.com/stw2/tokenizer-science-tax @ a77445db2440ec00d756a3346b3bddedc69e85eb
Reference checked 2026-09-14 13:15 UTC. Later commits, branches or plan changes do not retarget this attempt.
- Command
- python scripts/01_build_train_corpus.py --source sci_latex
- Working directory
- experiments/E09-vocabulary-allocation
- Configuration paths
- None
- Parameters
- quota 71,582,788 bytes; lines split on newline, trailing carriage return stripped, empty lines dropped, lines hard-split at 4,000 bytes; no randomness
- Environment
- 3.12.13; SentencePiece BPE trainer and Hugging Face tokenizers on CPU; no model weights; Apple silicon, 128 GB unified memory
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- proof-pile-2 arXiv train shard 0 · corpus source · M412 proof-pile-2 arXiv train shard 0 · dataset · Download
Delivered events
Registered
#1Builds the LaTeX training text from the first proof-pile-2 arXiv train shard. Unlogged; the source flag is inferred from the manifest order and file times.
Started
#2Started.
Succeeded
#3Completed with exit code 0.
Exit code 0.
- sci_latex.txt · Not redistributed: ask the Room owner, or rebuild with scripts/01_build_train_corpus.py at the pinned revisions · restricted · sha256 baf588da429a… · 72335954 bytes · M415
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.