Execution attempt · A72 · planned
Trains only the tied embedding matrix of the FOCUS-initialised transplant for 150M tokens and saves six checkpoints.
Pinned source and configuration
https://github.com/stw2/tokenizer-science-tax @ 80716b075735d909dacf90838f0e4ced6af46be8
Reference checked 2026-09-14 13:12 UTC. Later commits, branches or plan changes do not retarget this attempt.
- Command
- python scripts/07_train_stage2.py --model models/focus --data ../E05-control-arm/data/packed/apt4.bin --out runs/e7s2_focus --target-tokens 1.5e8
- Working directory
- experiments/E07-embedding-init
- Configuration paths
- None
- Parameters
- seq 2048; micro-batch 8; grad-accum 32 (524,288 tokens per step); lr 1e-4; warmup 50 steps, no decay; weight decay 0.1; clip 1.0; betas 0.9, 0.95; milestones 30,60,90,100,120,150M; trainable model.embed_tokens.weight only; stream data/packed/apt4.bin from offset 0, digest not recorded; the 150M checkpoint holds 150,470,656 tokens after 287 steps; times are the first and last checkpoint file times
- Environment
- not recorded; PyTorch CUDA, bf16, SDPA attention, non-reentrant gradient checkpointing, Liger fused cross-entropy; consumer CUDA GPU, 16 GB
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- models/focus · base model · M358 models/focus · raw output · Ask the reporter
Delivered events
Registered
#1Trains only the tied embedding matrix of the FOCUS-initialised transplant for 150M tokens and saves six checkpoints.
Started
#2Started.
Succeeded
#3Completed with exit code 0.
Exit code 0.
- ckpts_s2/focus/ckpt_00030M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 660e28a2ed2a… · 2722731597 bytes · M377
- ckpts_s2/focus/ckpt_00060M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 b80d81345321… · 2722731598 bytes · M378
- ckpts_s2/focus/ckpt_00090M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 7654b0613b9a… · 2722731598 bytes · M379
- ckpts_s2/focus/ckpt_00100M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 f6581063ca1b… · 2722731600 bytes · M380
- ckpts_s2/focus/ckpt_00120M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 c6400a87766b… · 2722731600 bytes · M381
- ckpts_s2/focus/ckpt_00150M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 9a1af1cb4a6d… · 2722731600 bytes · M382
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.