Finding

Sign in with GitHub
← Publications

Finding · P151 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B with 0%, 10% or 30% math and code, bits per byte is above the base model's on English web text, arXiv abstracts, LaTeX method sections, Python code and math_clean.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: All arms above base

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Bits per byteMetric.
texts
the English SlimPajama-6B holdout; English arXiv abstracts; English LaTeX method sections; English Python code; math_cleanTexts the statement covers.
count
5 textsNumber of texts covered.

Experimental provenance

Method and evaluation protocol
Comparison of the per-text endpoint values of the three arms.
Dataset
The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
Reported results
en: base 0.9336, 1.1068, 1.1131, 1.1092; sci_arxiv: base 0.7637, 0.9658, 0.9328, 0.8982; sci_latex: base 0.7218, 0.9150, 0.8774, 0.8522; sci_python: base 0.4046, 0.7161, 0.6039, 0.5580; math_clean: base 1.6722, 2.0022, 1.9309, 1.8480.
Uncertainty and replication
Point values without intervals. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

All arms above base

On each of the count texts, metric after budget tokens of recovery pretraining of subject is above base_model's value for every math and code share tested.

Key all_arms_above_base · version 34366aa6-1fab-4f7a-b4b8-62d099a44e13

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Exact references

derived from

22103b70-2086-49a6-aa07-bc8ac62683e3

Values on one text.

derived from

1abc95eb-7497-4ea0-bed9-5e443efdae89

Values on one text.

derived from

4e564231-f53d-4b71-a98c-4a192e92efcc

Values on one text.

derived from

10fa55f9-83be-40b9-8623-c04d980d89bb

Values on one text.

derived from

a6a9b812-aa29-4ba0-863f-98336d23c4f5

Values on one text.