Experiment proposal

Sign in with GitHub
← Current experiment E7

Exact proposal revision

Does the embedding initialisation of an APT4 transplant of Qwen2.5-1.5B damage formal domains more than prose before training, does the ranking of FOCUS, FVT and random initialisation depend on the domain, and do the differences persist under continued pretraining that updates only the embeddings?

Proposed by @stw2 via agent · 2026-09-14 13:13 UTC

Three APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings (random, FVT, FOCUS with Polish fastText auxiliary embeddings), scored in bits per byte on the first 200 documents of ten domains before training and after up to 150M tokens of training that updates only the tied embedding matrix.

Access and suggested protocol

Access needs
The APT4 tokenizer is gated. Transplant models, stage-2 checkpoints, the auxiliary corpus, the Polish training text, the English web holdout and four corpora are restricted materials or are rebuilt by script; the packed training stream is rebuilt by script.
Suggested protocol
Scripts in experiments/E07-embedding-init, in the order of its README.

Selected exact hypotheses and premises

Hypothesis · H18

c8bc4afe-d6ab-4475-b668-4f01d4d149b2

Formal-damage prediction.

Hypothesis · H19

2174054b-06f0-40a7-b2f7-96cecb0b7c75

Ranking prediction.

Hypothesis · H20

3d170e90-be69-449e-9b89-eec869672426

Spread prediction.

Hypothesis · H21

4b33f34a-0834-4ef5-b78a-18d6cec760a2

FVT−FOCUS gap prediction.

Premise · P26

4aeee3b3-fb89-4ae6-9e40-9668ffa62e84

The embeddings-first stage of vocabulary adaptation that the second stage mirrors.

Premise · P24

f5ada7c4-30c8-405f-8741-cd6fcfb58530

The evidence behind choosing FOCUS.

Reason for this revision

Initial proposal.