Finding

Sign in with GitHub
← Publications

Finding · P164 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-1.5B, the rescue fraction of a 30% math and code share on math_clean bits per byte is 0.5684 (95% interval 0.4996 to 0.6598); the gap to the base model is 0.2613 without math and code and 0.1128 with the 30% share.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Rescue fraction value

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-1.5BBase model of the transplant and reference of the gaps.
treatment value
30 percentMath and code share of the treatment.
control value
0 percentMath and code share of the control.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Bits per byteMetric of the gaps.
evaluation text
math_clean statementsText scored.
value
0.5684 ratioRescue fraction.
interval low
0.4996 ratioLower bound of the 95% bootstrap interval.
interval high
0.6598 ratioUpper bound of the 95% bootstrap interval.
control gap
0.2613 bits per byteGap of the 0% arm to the base model, positive meaning worse.
treatment gap
0.1128 bits per byteGap of the 30% arm to the base model, positive meaning worse.

Experimental provenance

Method and evaluation protocol
Per-document negative log-likelihood of the first 200 math_clean items under the base model and the final checkpoints of the 0% and 30% arms; bits per byte pooled over documents; R = 1 - (30% gap / 0% gap); 2,000 paired bootstrap resamples of documents, seed 20260707.
Dataset
The first 200 items of math_clean, an 800-item synthetic arithmetic set.Version: sha256 60ef593cf8e00bc45123c4cc776c52b5acd358393612d330bbb9348c4937acbc · Access: public
Reported results
R 0.5684 [0.4996, 0.6598]; 0% arm gap 0.2613 [0.2315, 0.2907]; 30% arm gap 0.1128.
Uncertainty and replication
95% percentile intervals from 2,000 paired bootstrap resamples. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Evidence references
analysis_conf.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/analysis_conf.jsonmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/metrics.jsonperdoc/t0_base_15b__math_clean.npz · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/perdoc/t0_base_15b__math_clean.npzperdoc/conf00_final__math_clean.npz · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/perdoc/conf00_final__math_clean.npzperdoc/conf30_final__math_clean.npz · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/perdoc/conf30_final__math_clean.npz
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. 1.5B model and a 1B-token budget. The confirmation analysis script was written after most confirmation measurements existed (retrospective plan version). Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Rescue fraction value

After budget tokens of recovery pretraining of subject, the rescue fraction of a math and code share of treatment_value against one of control_value, on metric (over evaluation_text where named), is value with its 95% bootstrap interval; control_gap and treatment_gap are the two arms' gaps to base_model, positive meaning worse.

Key rescue_fraction_value · version 8d38c22e-e0b7-4a12-bfb9-d0005b167458

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

math_clean statements

Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.

Key math_clean_statements · version 1c1e659f-9adf-49a7-8976-649eb293aaa7

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c