Finding

Sign in with GitHub
← Publications

Finding · P171 · Author-curated

Digit-probe accuracy of an APT4 FVT transplant of Qwen2.5-1.5B recovery-pretrained without math and code is 0.5689 after 200,278,016 tokens and 0.3967 after 1,000,341,504 tokens; Qwen2.5-1.5B scores 0.5856.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Values at checkpoints

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-1.5BBase model of the transplant and reference of the gaps.
procedure
Continued pretrainingTraining applied.
share
0 percentMath and code share of the run.
early tokens
200278016 tokensTraining tokens of the earlier checkpoint.
early value
0.5689 accuracyMetric at the earlier checkpoint.
final tokens
1000341504 tokensTraining tokens of the final checkpoint.
final value
0.3967 accuracyMetric at the final checkpoint.
base value
0.5856 accuracyMetric of the base model.

Experimental provenance

Method and evaluation protocol
Overall accuracy of 900 graded completions at the 200M-token and 1B-token checkpoints and for the base model.
Dataset
900 completions: 300 synthetic arithmetic items in three format cells, 4-shot, greedy; specification configs/probe_digits_spec.json.Version: sha256 e11c98fc129e4e4db6566140286c15a0fa6f88f7d7a4abdecad3a3b8f3ca53f8 · Access: public
Reported results
200M 0.5689; 1B 0.3967; base 0.5856.
Uncertainty and replication
One training run; no interval.
Limitations
One training run: checkpoint-to-checkpoint changes and evaluation noise are not separated. The transplant's generation configuration names end-of-sequence id 4 while training separated documents with id 2; answers are cut at the first newline.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Values at checkpoints

In one training run of subject under procedure, metric is early_value at early_tokens and final_value at final_tokens; base_value is the base model's.

Key values_at_checkpoints · version 0d8424f7-3e17-4327-b38b-6f375ac39213

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Continued pretraining

Further next-token-prediction training of a pretrained language model on additional text.

Key continued_pretraining · version f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c