Finding

Sign in with GitHub
← Publications

Finding · P157 · Author-curated

At the first checkpoint, after 100,139,008 tokens of recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B, bits per byte is 2.0426, 1.9824, 1.9466 on math_clean and 0.8468, 0.6936, 0.6641 on English Python code for the 0%, 10% and 30% math and code shares.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Values by share on two texts

subject
APT4 FVT transplantModel family recovered.
budget
100139008 tokensTraining tokens at the checkpoint.
metric
Bits per byteMetric.
text 1
math_clean statementsEvaluation text math_clean.
text 1 00
2.0426 bits per byteArm with 0% math and code on math_clean.
text 1 10
1.9824 bits per byteArm with 10% math and code on math_clean.
text 1 30
1.9466 bits per byteArm with 30% math and code on math_clean.
text 2
English Python codeEvaluation text English Python code.
text 2 00
0.8468 bits per byteArm with 0% math and code on English Python code.
text 2 10
0.6936 bits per byteArm with 10% math and code on English Python code.
text 2 30
0.6641 bits per byteArm with 30% math and code on English Python code.

Experimental provenance

Method and evaluation protocol
Bits per byte of each arm's first checkpoint row on the first 200 documents of math_clean and Python code.
Dataset
The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
Reported results
math_clean: 2.0426, 1.9824, 1.9466; sci_python: 0.8468, 0.6936, 0.6641.
Uncertainty and replication
Point values without intervals. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
Limitations
One training run per arm, so the intervals cover evaluation sampling only, not training variance. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Values by share on two texts

After budget tokens of recovery pretraining of subject, metric on text_1 is text_1_00, text_1_10 and text_1_30 and on text_2 is text_2_00, text_2_10 and text_2_30 for math and code shares of 0%, 10% and 30%.

Key values_by_share_on_two_texts · version 88d16551-a5e8-4c4d-91e9-503d03b1414e

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c

math_clean statements

Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.

Key math_clean_statements · version 1c1e659f-9adf-49a7-8976-649eb293aaa7

English Python code

331 Python files from GitHub, truncated at 700 words.

Key english_python_code · version 4acc878a-a4a3-45bd-bb03-2cd8541a7856