Finding

Sign in with GitHub
← Publications

Finding · P194 · Author-curated

The fall in bits per byte of the FVT-initialised APT4 transplant of Qwen2.5-1.5B after 100M tokens of embeddings-only training is between 0.707 (Polish PES examination questions) and 0.92 (English Python code) of its fall after 100M tokens of full-parameter training on the same token stream, over 9 domains.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Range over cells

quantity
fall in bits per byte from time zero after 100M tokens of embeddings-only training divided by the fall after 100M tokens of full-parameter training of the same FVT transplant on the same token streamQuantity ranged over.
model
APT4 transplant of Qwen2.5-1.5BModels measured.
cells
the nine primary domainsCells the range is taken over.
count cells
9 domainsCells.
min value
0.707 ratioSmallest value.
min cell
Polish PES examination questionsDomain with the smallest value.
max value
0.92 ratioLargest value.
max cell
English Python codeDomain with the largest value.

Experimental provenance

Method and evaluation protocol
Numerator: time-zero FVT row minus the 100M-token embeddings-only row. Denominator: the full-parameter experiment's time-zero FVT row minus its 100M-token row. Divided per domain.
Dataset
Bits-per-byte rows on the first 200 documents of the nine primary domains, from this experiment and from full-parameter continued pretraining of the same FVT transplant.Version: unspecified · Access: restricted
Reported results
sci_latex: 1.647 / 1.9168 = 0.859; sci_python: 2.1181 / 2.3017 = 0.92; math_clean: 0.7206 / 0.8726 = 0.826; sci_arxiv: 1.2647 / 1.5409 = 0.821; pl_wiki_sci: 1.6334 / 2.0411 = 0.8; pl_pes: 1.087 / 1.5369 = 0.707; pl: 1.4155 / 1.8038 = 0.785; en: 1.2675 / 1.5542 = 0.816; pl_informal: 0.9312 / 1.2172 = 0.765
Uncertainty and replication
Point estimates; the full-parameter rows have no per-document values, so no interval.
Limitations
The denominator comes from another experiment's scored rows (bpb_armB.jsonl). Descriptive endpoint without bins. One training run per initialisation at 1.5B. The trainer's loss and gradient-norm logs, the smoke-run gate results, the transfer digests and the token stream's digest were not kept. Windows of 2,048 tokens.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Values from results/analysis_s2.json, S2D3_embed_share_of_fullft.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Range over cells

The quantity ranges from min_value, in min_cell, to max_value, in max_cell, over count_cells cells.

Key range_over_cells · version abc309b2-cba4-4af7-a2b2-d0a31eb36b6d

APT4 transplant of Qwen2.5-1.5B

Qwen2.5-1.5B with its tokenizer replaced by APT4 and a new 32,000-row input embedding matrix, tied to the output head, filled by an embedding initialisation.

Key qwen_apt4_transplant · version c8bc4afe-d6ab-4475-b668-4f01d4d149b2

Exact references

related

4aeee3b3-fb89-4ae6-9e40-9668ffa62e84

The embeddings-first stage of vocabulary adaptation whose share of early recovery this measures.

related

131e534a-1b09-40c1-95fc-a4d5d63faae0

The full-parameter run of the same FVT transplant whose rows give the denominator.

related

f1355715-5d7f-4a04-a4bf-249f33aa288e

The full-parameter run of the same FVT transplant whose rows give the denominator.

related

b11964d8-81ba-4bf6-9364-302e0e5a677c

The full-parameter run of the same FVT transplant whose rows give the denominator.

related

ebd85d87-d358-44c8-aa85-5e8246d99b2f

The full-parameter run of the same FVT transplant whose rows give the denominator.

related

4a7dff3a-4d95-465e-a43e-bb0df933fc9e

The full-parameter run of the same FVT transplant whose rows give the denominator.

related

e2da0e41-7694-4b22-a1ea-14d13dc51f38

The full-parameter run of the same FVT transplant whose rows give the denominator.