Finding

Sign in with GitHub
← Publications

Finding · P150 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B with 0%, 10% or 30% math and code, bits per byte is below the base model's on each of four Polish texts: the FineWeb2-HQ holdout, PES examination questions, Wikipedia science articles and reviews.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: All arms below base

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Bits per byteMetric.
texts
the Polish FineWeb2-HQ holdout; Polish PES examination questions; Polish Wikipedia science articles; Polish reviewsTexts the statement covers.
count
4 textsNumber of texts covered.

Experimental provenance

Method and evaluation protocol
Comparison of the per-text endpoint values of the three arms.
Dataset
The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
Reported results
pl: base 1.4795, 0.9543, 0.9635, 0.9756; pl_pes: base 1.7108, 1.0524, 1.0679, 1.0758; pl_wiki_sci: base 1.3879, 0.9359, 0.9455, 0.9533; pl_informal: base 1.7548, 1.4965, 1.4947, 1.5059.
Uncertainty and replication
Point values without intervals. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

All arms below base

On each of the count texts, metric after budget tokens of recovery pretraining of subject is below base_model's value for every math and code share tested.

Key all_arms_below_base · version 31f58c70-7c36-46ac-89de-a49951f087b3

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Exact references

derived from

78c4f85d-fff2-45f1-8cab-cadefb6558a3

Values on one text.

derived from

09f49b47-ae82-4a35-bbfc-00e7bcce472f

Values on one text.

derived from

b9cc3c4f-7545-4c3a-9492-214ce5c47a5f

Values on one text.

derived from

a5f25f1e-7a1f-4666-966d-327bf336e9c4

Values on one text.