Finding

Sign in with GitHub
← Publications

Finding · P167 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant without math and code, the gap to the base model on math_clean bits per byte is 0.2613 with Qwen2.5-1.5B and 0.3305 with Qwen2.5-0.5B.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Values at two scales

subject
APT4 FVT transplantModel family recovered.
budget
1000000000 tokensRecovery pretraining tokens per arm.
quantity
gap to the base model after recovery pretraining without math and code, positive meaning worseQuantity compared.
metric
Bits per byteMetric.
evaluation text
math_clean statementsText scored.
larger model
Qwen2.5-1.5BLarger base model.
larger value
0.2613 bits per byteQuantity with the larger base model.
smaller model
Qwen2.5-0.5BSmaller base model.
smaller value
0.3305 bits per byteQuantity with the smaller base model.

Experimental provenance

Method and evaluation protocol
Values of the 0.5B and 1.5B analyses.
Dataset
The first 200 items of math_clean, an 800-item synthetic arithmetic set.Version: sha256 60ef593cf8e00bc45123c4cc776c52b5acd358393612d330bbb9348c4937acbc · Access: public
Reported results
1.5B 0.2613 [0.2315, 0.2907]; 0.5B 0.3305 [0.3129, 0.3492].
Uncertainty and replication
The 95% intervals of the two scales do not overlap; no test of the difference was computed. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
Two scales, one training run each; the 1.5B arms split the same tokens per step into micro-batch 8 and gradient accumulation 32. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Values at two scales

quantity on metric (over evaluation_text where named), after budget tokens of recovery pretraining of subject, is larger_value with larger_model as the base model and smaller_value with smaller_model.

Key scale_values · version 04457caa-cc84-4ccd-ac4d-bed74fdb2319

math_clean statements

Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.

Key math_clean_statements · version 1c1e659f-9adf-49a7-8976-649eb293aaa7

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Exact references

derived from

8d38c22e-e0b7-4a12-bfb9-d0005b167458

Values of one scale.

derived from

7eb877af-9ccc-4caa-ba55-b90d5716242e

Values of one scale.