Exact proposal revision
Does the embedding initialisation of an APT4 transplant of Qwen2.5-1.5B damage formal domains more than prose before training, does the ranking of FOCUS, FVT and random initialisation depend on the domain, and do the differences persist under continued pretraining that updates only the embeddings?
Three APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings (random, FVT, FOCUS with Polish fastText auxiliary embeddings), scored in bits per byte on the first 200 documents of ten domains before training and after up to 150M tokens of training that updates only the tied embedding matrix.
Access and suggested protocol
- Access needs
- The APT4 tokenizer is gated. Transplant models, stage-2 checkpoints, the auxiliary corpus, the Polish training text, the English web holdout and four corpora are restricted materials or are rebuilt by script; the packed training stream is rebuilt by script.
- Suggested protocol
- Scripts in experiments/E07-embedding-init, in the order of its README.
Selected exact hypotheses and premises
Premise · P26
4aeee3b3-fb89-4ae6-9e40-9668ffa62e84The embeddings-first stage of vocabulary adaptation that the second stage mirrors.
Reason for this revision
Initial proposal.