Finding

Sign in with GitHub
← Publications

Finding · P139 · Author-curated

During recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B with 0%, 10% or 30% math and code, bits per byte reaches 90% of its final recovery at 100,139,008 tokens on math_clean, Python code and the English holdout, at 200,278,016 tokens on the Polish holdout, and on Polish reviews at 200,278,016 tokens with the 0% share, 200,278,016 with 10% and 300,417,024 with 30%.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Recovery tokens by text

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant and reference of the gaps.
procedure
Continued pretrainingTraining applied.
measure
Tokens to 90% recoveryQuantity reported.
checkpoint interval
100000000 tokensTraining tokens between scored checkpoints.
math clean tokens
100139008 tokensMeasure on math_clean, equal in all three arms.
python code tokens
100139008 tokensMeasure on English Python code, equal in all three arms.
english holdout tokens
100139008 tokensMeasure on the English SlimPajama-6B holdout, equal in all three arms.
polish holdout tokens
200278016 tokensMeasure on the Polish FineWeb2-HQ holdout, equal in all three arms.
polish reviews tokens 0
200278016 tokensMeasure on Polish reviews without math and code.
polish reviews tokens 10
200278016 tokensMeasure on Polish reviews with the 10% share.
polish reviews tokens 30
300417024 tokensMeasure on Polish reviews with the 30% share.

Experimental provenance

Method and evaluation protocol
Bits per byte of each checkpoint row from the transplant at time zero to 1B tokens; the first checkpoint whose value has moved at least 90% of the way from the time-zero value to the last checkpoint's value.
Dataset
The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
Reported results
mix00: math_clean 100,139,008, sci_python 100,139,008, en 100,139,008, pl 200,278,016, pl_informal 200,278,016; mix10: math_clean 100,139,008, sci_python 100,139,008, en 100,139,008, pl 200,278,016, pl_informal 200,278,016; mix30: math_clean 100,139,008, sci_python 100,139,008, en 100,139,008, pl 200,278,016, pl_informal 300,417,024.
Uncertainty and replication
Resolution of one checkpoint interval (100M tokens); no interval estimate. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Evidence references
analysis.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/analysis.jsonmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/metrics.jsonbpb_mix00.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/bpb_mix00.jsonlbpb_mix10.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/bpb_mix10.jsonlbpb_mix30.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/2457520fc798b19bee6893ac58e0357ec4c467d5/experiments/E06-math-code-rescue/results/bpb_mix30.jsonl
Limitations
Every text reaches the threshold within the first three checkpoints, so the 100M-token spacing limits the comparison. The design reports this measure without a decision rule. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Recovery tokens by text

measure takes the stated token counts on the named texts in the training runs of subject under procedure; checkpoints are checkpoint_interval tokens apart.

Key recovery_tokens_by_text · version 59bfc2e8-6b8c-487d-8850-7c7c9e041de2

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

Tokens to 90% recovery

Training tokens at the first checkpoint where bits per byte on a text has moved at least 90% of the way from its value at the start of recovery pretraining to its value at the last checkpoint.

Key tokens_to_90pct_recovery · version bc2c559b-2e70-44d8-87ff-a07a94962628

Continued pretraining

Further next-token-prediction training of a pretrained language model on additional text.

Key continued_pretraining · version f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c