Finding

Sign in with GitHub
← Publications

Finding · P156 · Author-curated

Digit-probe accuracy of an APT4 FVT transplant of Qwen2.5-0.5B recovery-pretrained without math and code ranges from 0.1044 at 800,063,488 tokens to 0.3033 at 200,278,016 tokens over its ten checkpoints from 100M to 1B tokens.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Range over checkpoints

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant and reference of the gaps.
procedure
Continued pretrainingTraining applied.
share
0 percentMath and code share of the run.
minimum
0.1044 accuracyLowest value.
minimum tokens
800063488 tokensTraining tokens at the lowest value.
maximum
0.3033 accuracyHighest value.
maximum tokens
200278016 tokensTraining tokens at the highest value.
checkpoints
10 checkpointsCheckpoints compared, every 100M tokens.

Experimental provenance

Method and evaluation protocol
Overall accuracy of 900 graded completions at each of the ten checkpoints.
Dataset
900 completions: 300 synthetic arithmetic items in three format cells, 4-shot, greedy; specification configs/probe_digits_spec.json.Version: sha256 e11c98fc129e4e4db6566140286c15a0fa6f88f7d7a4abdecad3a3b8f3ca53f8 · Access: public
Reported results
minimum 0.1044 (800,063,488 tokens), maximum 0.3033 (200,278,016 tokens).
Uncertainty and replication
One training run; no interval.
Limitations
One training run: the spread mixes checkpoint-to-checkpoint changes with evaluation noise. The transplant's generation configuration names end-of-sequence id 4 while training separated documents with id 2; answers are cut at the first newline.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Range over checkpoints

Over the named checkpoints of one training run of subject under procedure, metric lies between minimum and maximum, reached at the stated training tokens.

Key range_over_checkpoints · version 6fe7b629-b8ab-4378-987c-7692f165a3d6

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Continued pretraining

Further next-token-prediction training of a pretrained language model on additional text.

Key continued_pretraining · version f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c