Finding

Sign in with GitHub
← Publications

Finding · P154 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B, digit-probe accuracy is 0.1233 without math and code, 0.1678 with a 10% share and 0.1644 with a 30% share; the base model scores 0.3967.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Values by share

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant and reference of the gaps.
procedure
Continued pretrainingTraining applied.
variable
Math and code shareQuantity varied.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Digit-probe accuracyMetric reported.
base value
0.3967 accuracyBase model.
value 00
0.1233 accuracyArm with 0% math and code.
value 10
0.1678 accuracyArm with 10% math and code.
value 30
0.1644 accuracyArm with 30% math and code.

Experimental provenance

Method and evaluation protocol
Overall accuracy of 900 graded completions per model at the 1B-token checkpoint and for the base model.
Dataset
900 completions: 300 synthetic arithmetic items in three format cells, 4-shot, greedy; specification configs/probe_digits_spec.json.Version: sha256 e11c98fc129e4e4db6566140286c15a0fa6f88f7d7a4abdecad3a3b8f3ca53f8 · Access: public
Reported results
0% 0.1233, 10% 0.1678, 30% 0.1644; base 0.3967.
Uncertainty and replication
No interval or test of the difference between the 10% and 30% arms was computed. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. Accuracy of one arm varies between its checkpoints by more than the difference between the 10% and 30% arms. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest. The transplant's generation configuration names end-of-sequence id 4 while training separated documents with id 2; answers are cut at the first newline.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Values by share

After the budget of procedure applied to subject, metric takes the stated value for each math and code share (roles ending in 00, 10 and 30) and base_value for the base model.

Key values_by_share · version 28d77606-f98b-43eb-9513-9f99a7935f24

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Continued pretraining

Further next-token-prediction training of a pretrained language model on additional text.

Key continued_pretraining · version f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Math and code share

Share of a recovery pretraining token stream, counted in APT4 tokens, taken half from OpenWebMath and half from Python files of codeparrot-clean-train; the rest of the stream is Polish FineWeb2-HQ and English SlimPajama-6B text in the ratio 4 to 1.

Key math_code_share · version f89e740e-7204-4148-8934-85fbc73f323c