Execution attempt

Sign in with GitHub
← Experiment E7 · Does the embedding initialisation of an APT4 transplant of Qwen2.5-1.5B damage formal domains more than prose before training, does the ranking of FOCUS, FVT and random initialisation depend on the domain, and do the differences persist under continued pretraining that updates only the embeddings?

Execution attempt · A71 · planned

Trains only the tied embedding matrix of the FVT-initialised transplant for 150M tokens and saves six checkpoints.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 80716b075735d909dacf90838f0e4ced6af46be8

Reference checked 2026-09-14 13:12 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/07_train_stage2.py --model models/fvt --data ../E05-control-arm/data/packed/apt4.bin --out runs/e7s2_fvt --target-tokens 1.5e8
Working directory
experiments/E07-embedding-init
Configuration paths
None
Parameters
seq 2048; micro-batch 8; grad-accum 32 (524,288 tokens per step); lr 1e-4; warmup 50 steps, no decay; weight decay 0.1; clip 1.0; betas 0.9, 0.95; milestones 30,60,90,100,120,150M; trainable model.embed_tokens.weight only; stream data/packed/apt4.bin from offset 0, digest not recorded; the 150M checkpoint holds 150,470,656 tokens after 287 steps; times are the first and last checkpoint file times
Environment
not recorded; PyTorch CUDA, bf16, SDPA attention, non-reentrant gradient checkpointing, Liger fused cross-entropy; consumer CUDA GPU, 16 GB
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • models/fvt · base model · M357 models/fvt · raw output · Ask the reporter

Delivered events

  1. Registered

    #1

    Trains only the tied embedding matrix of the FVT-initialised transplant for 150M tokens and saves six checkpoints.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • ckpts_s2/fvt/ckpt_00030M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 88ee0f782de9… · 2722731593 bytes · M371
    • ckpts_s2/fvt/ckpt_00060M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 ba70740820b3… · 2722731594 bytes · M372
    • ckpts_s2/fvt/ckpt_00090M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 4bf6d9120b67… · 2722731594 bytes · M373
    • ckpts_s2/fvt/ckpt_00100M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 3cffb276972c… · 2722731596 bytes · M374
    • ckpts_s2/fvt/ckpt_00120M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 c33d118f8a14… · 2722731596 bytes · M375
    • ckpts_s2/fvt/ckpt_00150M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 98c07fc06c15… · 2722731596 bytes · M376

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.