Execution attempt

Sign in with GitHub
← Experiment E7 · Does the embedding initialisation of an APT4 transplant of Qwen2.5-1.5B damage formal domains more than prose before training, does the ranking of FOCUS, FVT and random initialisation depend on the domain, and do the differences persist under continued pretraining that updates only the embeddings?

Execution attempt · A73 · planned

Trains only the tied embedding matrix of the random-initialised transplant for 150M tokens and saves six checkpoints.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ 80716b075735d909dacf90838f0e4ced6af46be8

Reference checked 2026-09-14 13:12 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/07_train_stage2.py --model models/random --data ../E05-control-arm/data/packed/apt4.bin --out runs/e7s2_random --target-tokens 1.5e8
Working directory
experiments/E07-embedding-init
Configuration paths
None
Parameters
seq 2048; micro-batch 8; grad-accum 32 (524,288 tokens per step); lr 1e-4; warmup 50 steps, no decay; weight decay 0.1; clip 1.0; betas 0.9, 0.95; milestones 30,60,90,100,120,150M; trainable model.embed_tokens.weight only; stream data/packed/apt4.bin from offset 0, digest not recorded; the 150M checkpoint holds 150,470,656 tokens after 287 steps; times are the first and last checkpoint file times
Environment
not recorded; PyTorch CUDA, bf16, SDPA attention, non-reentrant gradient checkpointing, Liger fused cross-entropy; consumer CUDA GPU, 16 GB
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • models/random · base model · M356 models/random · raw output · Ask the reporter

Delivered events

  1. Registered

    #1

    Trains only the tied embedding matrix of the random-initialised transplant for 150M tokens and saves six checkpoints.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • ckpts_s2/random/ckpt_00030M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 e0eb17327149… · 2722731599 bytes · M383
    • ckpts_s2/random/ckpt_00060M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 88315aa7a319… · 2722731600 bytes · M384
    • ckpts_s2/random/ckpt_00090M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 23013fb8e25e… · 2722731600 bytes · M385
    • ckpts_s2/random/ckpt_00100M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 6ecbc4ded4ab… · 2722731602 bytes · M386
    • ckpts_s2/random/ckpt_00120M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 3a7df5039087… · 2722731602 bytes · M387
    • ckpts_s2/random/ckpt_00150M · Not redistributed: ask the Room owner, or rebuild with scripts/07_train_stage2.py at the pinned revisions. The tokenizer files are saved with a Qwen2 tokenizer class; load them through load_tokenizer() in scripts/03_score_bpb.py · restricted · sha256 87bbf38dd8e9… · 2722731602 bytes · M388

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.