- Method and evaluation protocol
- Bits per byte of each checkpoint row from the transplant at time zero to 1B tokens; the first checkpoint whose value has moved at least 90% of the way from the time-zero value to the last checkpoint's value.
- Dataset
- The first 200 documents of ten evaluation texts: Polish and English web holdouts, English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, Polish Wikipedia science articles, PES questions and reviews, math_clean.Version: unspecified · Access: restricted
- Reported results
- mix00: math_clean 100,139,008, sci_python 100,139,008, en 100,139,008, pl 200,278,016, pl_informal 200,278,016; mix10: math_clean 100,139,008, sci_python 100,139,008, en 100,139,008, pl 200,278,016, pl_informal 200,278,016; mix30: math_clean 100,139,008, sci_python 100,139,008, en 100,139,008, pl 200,278,016, pl_informal 300,417,024.
- Uncertainty and replication
- Resolution of one checkpoint interval (100M tokens); no interval estimate. One training run per arm, so the intervals cover evaluation sampling only, not training variance.
- Limitations
- Every text reaches the threshold within the first three checkpoints, so the 100M-token spacing limits the comparison. The design reports this measure without a decision rule. The OpenWebMath source of the math and code share was streamed without a recorded revision, so the training streams with math and code cannot be rebuilt byte-identically; the packed streams are identified by digest.