Finding

Sign in with GitHub
← Publications

Finding · P153 · Author-curated

After 1B tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B, the rescue fraction of a 30% math and code share on Python code bits per byte is 0.5074 (point estimate).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Rescue fraction point estimate

subject
APT4 FVT transplantModel family recovered.
base model
Qwen2.5-0.5BBase model of the transplant and reference of the gaps.
procedure
Continued pretrainingTraining applied.
variable
Math and code shareQuantity varied.
treatment value
30 percentMath and code share of the treatment.
control value
0 percentMath and code share of the control.
budget
1000000000 tokensRecovery pretraining tokens per arm.
metric
Bits per byteMetric of the gaps.
evaluation text
English Python codeText scored.
value
0.5074 ratioRescue fraction.

Experimental provenance

Method and evaluation protocol
R = 1 - (30% arm minus base) / (0% arm minus base) from the 1B-token checkpoint rows and the base model's row on the first 200 Python code documents.
Dataset
The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
Reported results
R 0.5074; base 0.4046, 0% arm 0.7161, 30% arm 0.5580.
Uncertainty and replication
Point estimate without an interval. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. Python code is not an endpoint of the rescue hypothesis. Dataset and base-model revisions were not recorded in the original run; the pinned revisions are inferred. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Rescue fraction point estimate

After the budget of procedure applied to subject, the treatment's rescue fraction against the control on metric over evaluation_text is value, without an interval.

Key rescue_fraction_point · version 7b8c5b9c-15bf-4f49-b994-ef0e1f6f63ba

Qwen2.5-0.5B

The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

Key qwen2_5_0_5b · version f89e740e-7204-4148-8934-85fbc73f323c

English Python code

331 Python files from GitHub, truncated at 700 words.

Key english_python_code · version 4acc878a-a4a3-45bd-bb03-2cd8541a7856

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

Continued pretraining

Further next-token-prediction training of a pretrained language model on additional text.

Key continued_pretraining · version f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

Math and code share

Share of a recovery pretraining token stream, counted in APT4 tokens, taken half from OpenWebMath and half from Python files of codeparrot-clean-train; the rest of the stream is Polish FineWeb2-HQ and English SlimPajama-6B text in the ratio 4 to 1.

Key math_code_share · version f89e740e-7204-4148-8934-85fbc73f323c