Finding

Sign in with GitHub
← Publications

Finding · P165 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-1.5B, the rescue fraction of a 30% math and code share on digit-probe accuracy is 0.4706 (95% interval 0.3571 to 0.5871); the gap to the base model is 0.1889 without math and code and 0.1000 with the 30% share.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Rescue fraction value

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-1.5BBase model of the transplant and reference of the gaps.
treatment value
30 percentMath and code share of the treatment.
control value
0 percentMath and code share of the control.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Digit-probe accuracyMetric of the gaps.
value
0.4706 ratioRescue fraction.
interval low
0.3571 ratioLower bound of the 95% bootstrap interval.
interval high
0.5871 ratioUpper bound of the 95% bootstrap interval.
control gap
0.1889 accuracyGap of the 0% arm to the base model, positive meaning worse.
treatment gap
0.1000 accuracyGap of the 30% arm to the base model, positive meaning worse.

Experimental provenance

Method and evaluation protocol
Greedy 4-shot completions of 300 synthetic arithmetic items in three format cells (900 per model) by the base model and the final checkpoints of the 0% and 30% arms, graded numerically; accuracy gaps signed so that positive means worse; R = 1 - (30% gap / 0% gap); 2,000 paired bootstrap resamples of items, seed 20260707.
Dataset
900 completions: 300 synthetic arithmetic items in three format cells, 4-shot, greedy; specification configs/probe_digits_spec.json.Version: sha256 e11c98fc129e4e4db6566140286c15a0fa6f88f7d7a4abdecad3a3b8f3ca53f8 · Access: public
Reported results
R 0.4706 [0.3571, 0.5871]; 0% arm gap 0.1889 [0.1567, 0.2211]; 30% arm gap 0.1000. Accuracy: base model 0.5856, 0% arm 0.3967, 30% arm 0.4856.
Uncertainty and replication
95% percentile intervals from 2,000 paired bootstrap resamples. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Evidence references
analysis_conf.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/analysis_conf.jsonmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/metrics.jsonprobe_digits_raw/conf00/t0_base.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/probe_digits_raw/conf00/t0_base.jsonlprobe_digits_raw/conf00/ckpt_01000M.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/probe_digits_raw/conf00/ckpt_01000M.jsonlprobe_digits_raw/conf30/ckpt_01000M.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/probe_digits_raw/conf30/ckpt_01000M.jsonl
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. The 95% interval of the rescue fraction crosses 0.5. Accuracy of the 0% arm varies between its checkpoints (0.5689 at 200M tokens). The confirmation analysis script was written after most confirmation measurements existed (retrospective plan version). Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest. The transplant's generation configuration names end-of-sequence id 4 while training separated documents with id 2; answers are cut at the first newline.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Rescue fraction value

After budget tokens of recovery pretraining of subject, the rescue fraction of a math and code share of treatment_value against one of control_value, on metric (over evaluation_text where named), is value with its 95% bootstrap interval; control_gap and treatment_gap are the two arms' gaps to base_model, positive meaning worse.

Key rescue_fraction_value · version 8d38c22e-e0b7-4a12-bfb9-d0005b167458

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c