Finding

Sign in with GitHub
← Publications

Finding · P145 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B, bits per byte on English arXiv abstracts is 0.9658 without math and code, 0.9328 with a 10% share and 0.8982 with a 30% share; Qwen2.5-0.5B scores 0.7637 and the transplant before recovery pretraining 2.6268.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Endpoint values

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Bits per byteMetric.
evaluation text
English arXiv abstractsText scored.
base value
0.7637 bits per byteBase model.
start value
2.6268 bits per byteTransplant before recovery pretraining.
value 00
0.9658 bits per byteArm with 0% math and code.
value 10
0.9328 bits per byteArm with 10% math and code.
value 30
0.8982 bits per byteArm with 30% math and code.

Experimental provenance

Method and evaluation protocol
Bits per byte of the base model's and the transplant's time-zero rows and of each arm's 1B-token checkpoint row on the first 200 documents of the text; windows of 2,048 tokens.
Dataset
The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
Reported results
base 0.7637; transplant at start 2.6268; 0% 0.9658; 10% 0.9328; 30% 0.8982.
Uncertainty and replication
Point values without intervals. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. Bits per byte compares models with different tokenizers on the same bytes. Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Endpoint values

On evaluation_text, metric after budget tokens of recovery pretraining of subject is value_00, value_10 and value_30 for math and code shares of 0%, 10% and 30%; base_value is base_model's value and start_value is subject's value before recovery pretraining.

Key endpoint_values · version 78c4f85d-fff2-45f1-8cab-cadefb6558a3

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

English arXiv abstracts

784 abstracts of arXiv papers with January 2024 identifiers, at least 40 words each.

Key english_arxiv_abstracts · version 490c5f75-43e6-4d68-ba35-6a0097992182

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c