Finding

Sign in with GitHub
← Publications

Finding · P137 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B, the rescue fraction of a 30% math and code share on digit-probe accuracy is 0.1504 (95% interval 0.0931 to 0.2060); the gap to the base model is 0.2733 without math and code and 0.2322 with the 30% share.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Rescue fraction value

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant and reference of the gaps.
treatment value
30 percentMath and code share of the treatment.
control value
0 percentMath and code share of the control.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Digit-probe accuracyMetric of the gaps.
value
0.1504 ratioRescue fraction.
interval low
0.0931 ratioLower bound of the 95% bootstrap interval.
interval high
0.2060 ratioUpper bound of the 95% bootstrap interval.
control gap
0.2733 accuracyGap of the 0% arm to the base model, positive meaning worse.
treatment gap
0.2322 accuracyGap of the 30% arm to the base model, positive meaning worse.

Experimental provenance

Method and evaluation protocol
Greedy 4-shot completions of 300 synthetic arithmetic items in three format cells (900 per model) by the base model and the final checkpoints of the 0% and 30% arms, graded numerically; accuracy gaps signed so that positive means worse; R = 1 - (30% gap / 0% gap); 2,000 paired bootstrap resamples of items, seed 20260707.
Dataset
900 completions: 300 synthetic arithmetic items in three format cells, 4-shot, greedy; specification configs/probe_digits_spec.json.Version: sha256 e11c98fc129e4e4db6566140286c15a0fa6f88f7d7a4abdecad3a3b8f3ca53f8 · Access: public
Reported results
R 0.1504 [0.0931, 0.2060]; 0% arm gap 0.2733 [0.2411, 0.3033]; 30% arm gap 0.2322. Accuracy: base model 0.3967, 0% arm 0.1233, 30% arm 0.1644.
Uncertainty and replication
95% percentile intervals from 2,000 paired bootstrap resamples. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Evidence references
analysis.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/analysis.jsonmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/metrics.jsonprobe_digits_raw/mix00/t0_base.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/probe_digits_raw/mix00/t0_base.jsonlprobe_digits_raw/mix00/ckpt_01000M.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/probe_digits_raw/mix00/ckpt_01000M.jsonlprobe_digits_raw/mix30/ckpt_01000M.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/probe_digits_raw/mix30/ckpt_01000M.jsonl
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. Accuracy of one arm varies between its checkpoints (the 0% arm from 0.1044 to 0.3033), which the item bootstrap does not cover. The design's Holm adjustment and no-dose-effect clause were not computed. Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest. The transplant's generation configuration names end-of-sequence id 4 while training separated documents with id 2; answers are cut at the first newline.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Rescue fraction value

After budget tokens of recovery pretraining of subject, the rescue fraction of a math and code share of treatment_value against one of control_value, on metric (over evaluation_text where named), is value with its 95% bootstrap interval; control_gap and treatment_gap are the two arms' gaps to base_model, positive meaning worse.

Key rescue_fraction_value · version 8d38c22e-e0b7-4a12-bfb9-d0005b167458

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Exact references

related

d664a095-2689-445e-ac40-a85c6202a460

Tests one composition of the adaptation data that the claim leaves unspecified.