Finding

Sign in with GitHub
← Publications

Finding · P125 · Author-curated

On the English SlimPajama holdout, the bits per byte of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer exceed those of Qwen2.5-1.5B with its own tokenizer by 0.2371, 0.1849, 0.1594, 0.1442 and 0.1328 after 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Gap trajectory

metric
Bits per byteQuantity whose difference is taken.
subject
Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continuedModel whose metric is reduced by the comparator's.
comparator
Qwen2.5-1.5B continued with its own tokenizerModel whose metric is subtracted.
evaluation text
English SlimPajama holdoutText scored.
language
EnglishLanguage of the text.
milestones
100,139,008; 200,278,016; 300,417,024; 400,031,744; 500,170,752 tokensTraining tokens at which the gaps are measured.
gaps
+0.2371; +0.1849; +0.1594; +0.1442; +0.1328Subject minus comparator at each milestone, in bits per byte.
setting
the same document sequence, constant learning rate before any decay anneal, one run per armTraining both models received.

Experimental provenance

Method and evaluation protocol
Next-token negative log-likelihood in bits summed over non-overlapping 2048-token windows of each document, the first token of each window unscored, divided by the documents' UTF-8 bytes; bf16 on Apple silicon, 128 GB unified memory; arm-B tokenizers loaded through a whitespace canary and the frozen APT4 reference.
Dataset
The first 200 of 2,000 held-out SlimPajama-6B documents.Version: unspecified · Access: restricted
Reported results
Gaps +0.2371, +0.1849, +0.1594, +0.1442, +0.1328; successive ratios 0.7799, 0.8621, 0.9042, 0.9209.
Uncertainty and replication
One run per arm and one scoring pass per milestone; no interval. Five points do not distinguish decay toward zero from decay toward a positive residual.
Limitations
One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept. The extension to 1B tokens that was to decide between a permanent residual and slow convergence was never launched.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Gap trajectory

The subject's metric minus the comparator's metric takes the listed values at the listed training-token milestones.

Key gap_trajectory · version 131e534a-1b09-40c1-95fc-a4d5d63faae0

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

English SlimPajama holdout

The first 200 of 2,000 SlimPajama-6B documents held out from a seed-42 shuffled stream and never trained on.

Key slimpajama_english_holdout · version 0c90c839-74ad-4f13-8f54-0b2dfbc9afbb

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584