Finding

Sign in with GitHub
← Publications

Finding · P168 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant without math and code, the gap to the base model on digit-probe accuracy is 0.1889 with Qwen2.5-1.5B and 0.2733 with Qwen2.5-0.5B.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Values at two scales

subject
APT4 FVT transplantModel family recovered.
budget
1000000000 tokensRecovery pretraining tokens per arm.
quantity
gap to the base model after recovery pretraining without math and code, positive meaning worseQuantity compared.
larger model
Qwen2.5-1.5BLarger base model.
larger value
0.1889 accuracyQuantity with the larger base model.
smaller model
Qwen2.5-0.5BSmaller base model.
smaller value
0.2733 accuracyQuantity with the smaller base model.

Experimental provenance

Method and evaluation protocol
Values of the 0.5B and 1.5B analyses.
Dataset
900 completions: 300 synthetic arithmetic items in three format cells, 4-shot, greedy; specification configs/probe_digits_spec.json.Version: sha256 e11c98fc129e4e4db6566140286c15a0fa6f88f7d7a4abdecad3a3b8f3ca53f8 · Access: public
Reported results
1.5B 0.1889 [0.1567, 0.2211]; 0.5B 0.2733 [0.2411, 0.3033].
Uncertainty and replication
The 95% intervals of the two scales do not overlap; no test of the difference was computed. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
Two scales, one training run each; the 1.5B arms split the same tokens per step into micro-batch 8 and gradient accumulation 32. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest. The transplant's generation configuration names end-of-sequence id 4 while training separated documents with id 2; answers are cut at the first newline.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Values at two scales

quantity on metric (over evaluation_text where named), after budget tokens of recovery pretraining of subject, is larger_value with larger_model as the base model and smaller_value with smaller_model.

Key scale_values · version 04457caa-cc84-4ccd-ac4d-bed74fdb2319

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Exact references

derived from

d4a72d5f-d463-4ccf-9284-4197271f39bf

Values of one scale.

derived from

09b0327f-40c2-43fb-b6a9-9bbb6fa432da

Values of one scale.