Experiment · E7
Does the embedding initialisation of an APT4 transplant of Qwen2.5-1.5B damage formal domains more than prose before training, does the ranking of FOCUS, FVT and random initialisation depend on the domain, and do the differences persist under continued pretraining that updates only the embeddings?
Three APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings (random, FVT, FOCUS with Polish fastText auxiliary embeddings), scored in bits per byte on the first 200 documents of ten domains before training and after up to 150M tokens of training that updates only the tied embedding matrix.
Prerequisites and protocol
- Access needs
- The APT4 tokenizer is gated. Transplant models, stage-2 checkpoints, the auxiliary corpus, the Polish training text, the English web holdout and four corpora are restricted materials or are rebuilt by script; the packed training stream is rebuilt by script.
- Suggested protocol
- Scripts in experiments/E07-embedding-init, in the order of its README.
Accepted plan
Plan accepted by @stw2
Stage 2: 150M tokens of embeddings-only continued pretraining of the three transplants, with bits-per-byte recovery at six checkpoints- requires models/fvt · base model · M357 models/fvt · Ask the reporter
- requires models/focus · base model · M358 models/focus · Ask the reporter
- requires models/random · base model · M356 models/random · Ask the reporter
- requires APT4 tokenizer · reference tokenizer · M3 APT4 tokenizer (Bielik-PL-11B-v3.0-Instruct) · Ask the reporter
- requires bpb_e7.jsonl · time zero rows · M367 bpb_e7.jsonl · Download
- requires perdoc/ · time zero rows · M368 perdoc/ · Download
- requires bpb_armB.jsonl · reference values · M169 bpb_armB.jsonl as originally written · Rebuild via attempt
- requires pl: pl_eval.jsonl · evaluation data · M129 pl_eval.jsonl · Download
- requires en: en_eval.jsonl · evaluation data · M130 en_eval.jsonl · Ask the reporter
- requires pl_informal: pl-informal.jsonl · evaluation data · M33 pl-informal.jsonl · Ask the reporter
- requires pl_wiki_sci: pl-wiki-science.jsonl · evaluation data · M32 pl-wiki-science.jsonl · Download
- requires pl_pes: pl-science-pes.jsonl · evaluation data · M31 pl-science-pes.jsonl · Ask the reporter
- requires sci_arxiv: en-arxiv-abstracts.jsonl · evaluation data · M27 en-arxiv-abstracts.jsonl · Download
- requires sci_latex: en-latex-methods.jsonl · evaluation data · M28 en-latex-methods.jsonl · Ask the reporter
- requires sci_python: en-python-code.jsonl · evaluation data · M29 en-python-code.jsonl · Ask the reporter
- requires sci_gsm8k: en-gsm8k.jsonl · evaluation data · M30 en-gsm8k.jsonl · Download
- requires math_clean: math_clean.jsonl · evaluation data · M128 math_clean evaluation statements · Download
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 14
14 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
planned attempt · stw2/tokenizer-science-tax @ e54c7e02bb1e · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ e54c7e02bb1e · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ e54c7e02bb1e · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ e54c7e02bb1e · under an accepted plan
A67 Writes flat headline values from the stage-1 analysis, the build statistics and the auxiliary manifest.
Succeededplanned attempt · stw2/tokenizer-science-tax @ e54c7e02bb1e · under an accepted plan
A68 Re-scores the base model and the three unchanged transplants with whitespace-gated tokenizer loads after the erratum.
Succeededrerun attempt · stw2/tokenizer-science-tax @ 6c8a461e502d · under an accepted plan
A69 Recomputes the stage-1 endpoints, bins and gates from the re-scored rows with the unchanged analysis script.
Succeededrerun attempt · stw2/tokenizer-science-tax @ 6c8a461e502d · under an accepted plan
rerun attempt · stw2/tokenizer-science-tax @ 6c8a461e502d · under an accepted plan
A71 Trains only the tied embedding matrix of the FVT-initialised transplant for 150M tokens and saves six checkpoints.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 80716b075735 · under an accepted plan
A72 Trains only the tied embedding matrix of the FOCUS-initialised transplant for 150M tokens and saves six checkpoints.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 80716b075735 · under an accepted plan
A73 Trains only the tied embedding matrix of the random-initialised transplant for 150M tokens and saves six checkpoints.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 80716b075735 · under an accepted plan
A74 Scores the 18 stage-2 checkpoints in bits per byte on ten domains with whitespace-gated tokenizer loads.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 80716b075735 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 82a236f723c0 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 82a236f723c0 · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Embedding initialisation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →