Finding

Sign in with GitHub
← Publications

Finding · P162 · Author-curated

Bits per byte aggregated from per-document scores of the final checkpoints of three recovery pretraining runs of an APT4 FVT transplant of Qwen2.5-0.5B differs from their batched checkpoint scores by at most 0.0015 on math_clean and the Polish holdout.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Scorer difference at most

subject
APT4 FVT transplantModels scored.
metric
Bits per byteMetric.
scorer 1
per-document scorer, one window per forward passFirst scorer.
scorer 2
checkpoint scorer, windows padded into batches of 8Second scorer.
texts
math_clean; Polish FineWeb2-HQ holdoutTexts compared.
value
0.0015 bits per byteLargest absolute difference.
band
0.005 bits per byteTolerance set before measurement.

Experimental provenance

Method and evaluation protocol
Absolute differences between aggregated per-document bits per byte and the 1B-token checkpoint rows, per arm and text.
Dataset
The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
Reported results
largest difference 0.001506 (30% arm, math_clean); tolerance 0.005.
Uncertainty and replication
Deterministic scoring; no interval.
Limitations
Checked on two texts only; the differences come from bf16 numerics of batched and unbatched forward passes. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Scorer difference at most

The metric of each scored model on each of the texts differs between scorer_1 and scorer_2 by at most value, inside the tolerance band.

Key scorer_difference_at_most · version ea679d70-4f6a-49fb-ba27-b0083b779395

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

APT4 FVT transplant

A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c