Finding

Sign in with GitHub
← Publications

Finding · P136 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B, the rescue fraction of a 30% math and code share on math_clean bits per byte is 0.4664 (95% interval 0.4324 to 0.4981); the gap to the base model is 0.3305 without math and code and 0.1764 with the 30% share.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Rescue fraction value

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant and reference of the gaps.
treatment value
30 percentMath and code share of the treatment.
control value
0 percentMath and code share of the control.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Bits per byteMetric of the gaps.
evaluation text
math_clean statementsText scored.
value
0.4664 ratioRescue fraction.
interval low
0.4324 ratioLower bound of the 95% bootstrap interval.
interval high
0.4981 ratioUpper bound of the 95% bootstrap interval.
control gap
0.3305 bits per byteGap of the 0% arm to the base model, positive meaning worse.
treatment gap
0.1764 bits per byteGap of the 30% arm to the base model, positive meaning worse.

Experimental provenance

Method and evaluation protocol
Per-document negative log-likelihood of the first 200 math_clean items under the base model and the final checkpoints of the 0% and 30% arms; bits per byte pooled over documents; R = 1 - (30% gap / 0% gap); 2,000 paired bootstrap resamples of documents, seed 20260707.
Dataset
The first 200 items of math_clean, an 800-item synthetic arithmetic set.Version: sha256 60ef593cf8e00bc45123c4cc776c52b5acd358393612d330bbb9348c4937acbc · Access: public
Reported results
R 0.4664 [0.4324, 0.4981]; 0% arm gap 0.3305 [0.3129, 0.3492]; 30% arm gap 0.1764.
Uncertainty and replication
95% percentile intervals from 2,000 paired bootstrap resamples. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Evidence references
analysis.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/analysis.jsonmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/metrics.jsonperdoc/t0_base__math_clean.npz · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/perdoc/t0_base__math_clean.npzperdoc/mix00_final__math_clean.npz · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/perdoc/mix00_final__math_clean.npzperdoc/mix30_final__math_clean.npz · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/perdoc/mix30_final__math_clean.npz
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. 0.5B model and a 1B-token budget. Per-document aggregates differ from the checkpoint rows by up to 0.0015 bits per byte (batching numerics). The design's Holm adjustment and no-dose-effect clause were not computed. Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Rescue fraction value

After budget tokens of recovery pretraining of subject, the rescue fraction of a math and code share of treatment_value against one of control_value, on metric (over evaluation_text where named), is value with its 95% bootstrap interval; control_gap and treatment_gap are the two arms' gaps to base_model, positive meaning worse.

Key rescue_fraction_value · version 8d38c22e-e0b7-4a12-bfb9-d0005b167458

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

math_clean statements

Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.

Key math_clean_statements · version 1c1e659f-9adf-49a7-8976-649eb293aaa7

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Exact references

related

d664a095-2689-445e-ac40-a85c6202a460

Tests one composition of the adaptation data that the claim leaves unspecified.