Finding

Sign in with GitHub
← Publications

Finding · P182 · Author-curated

After 150M tokens of embeddings-only training, bits per byte of the random-initialised APT4 transplant of Qwen2.5-1.5B exceeds that of the FVT-initialised transplant by 1.4148 bits per byte on English Python code (95% interval 1.374 to 1.4596).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Difference with interval

metric
Bits per byteQuantity compared.
model
APT4 transplant of Qwen2.5-1.5BModels measured.
subject
Random embedding initialisationInitialisation of the model whose value is the minuend.
comparator
Frequency-based Vocabulary TransferInitialisation of the model whose value is subtracted.
evaluation text
English Python codeCorpus scored, first 200 documents.
training tokens
150000000 tokensTokens of embeddings-only continued pretraining, nominal checkpoint size, before the value is measured.
value
1.4148 bits per byteDifference.
interval low
1.374 bits per byteLower bound of the 95% interval.
interval high
1.4596 bits per byteUpper bound of the 95% interval.
bin
RESIDUALPre-registered class: CLOSED, PARTIAL or RESIDUAL.

Experimental provenance

Method and evaluation protocol
Paired document bootstrap (1000 resamples, seed 20260705) of random-minus-FVT bits per byte at the 150M-token checkpoints.
Dataset
First 200 documents of the evaluation domains, scored at the 150M-token checkpoints of embeddings-only training on the packed APT4 stream.Version: unspecified · Access: restricted
Reported results
1.4148 [1.374, 1.4596], RESIDUAL; random 2.3822, FVT 0.9674.
Uncertainty and replication
95% paired document bootstrap interval.
Limitations
One training run per initialisation at 1.5B. The trainer's loss and gradient-norm logs, the smoke-run gate results, the transfer digests and the token stream's digest were not kept. Windows of 2,048 tokens.

Author’s note

Values from results/analysis_s2.json, S2P_random_floor.sci_python and S2D1_curves.<arm>.ckpt_00150M.sci_python.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Difference with interval

subject's metric minus comparator's metric on evaluation_text after training_tokens is value, with a 95% paired bootstrap interval from interval_low to interval_high; bin is the pre-registered class of the value.

Key difference_with_interval · version bea59b66-839e-44b1-99d3-10123fbb043a

Frequency-based Vocabulary Transfer

An embedding initialisation for a replaced tokenizer's vocabulary.

Key fvt_init · version 01087ca0-917c-46f8-a25e-ebc786018d72

English Python code

331 Python files from GitHub, truncated at 700 words.

Key english_python_code · version 4acc878a-a4a3-45bd-bb03-2cd8541a7856

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 transplant of Qwen2.5-1.5B

Qwen2.5-1.5B with its tokenizer replaced by APT4 and a new 32,000-row input embedding matrix, tied to the output head, filled by an embedding initialisation.

Key qwen_apt4_transplant · version c8bc4afe-d6ab-4475-b668-4f01d4d149b2

Random embedding initialisation

An embedding initialisation for a replaced tokenizer's vocabulary that samples new vectors at random.

Key random_init · version cae5a2ea-f055-4e01-bb4b-86bb0b6c0fd3

Exact references

extends

87d352dc-dabc-4bcc-8c74-2bf51e2e548d

Follows the random and FVT transplants on a formal domain through embeddings-only training.