Finding

Sign in with GitHub
← Publications

Finding · P112 · Author-curated

After 0.5B tokens of the same continued pretraining, bits per byte on math_clean arithmetic statements is 1.7672 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 1.5157 for Qwen2.5-1.5B with its own tokenizer.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Bits per byteQuantity compared.
subject
Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continuedModel whose value is subject_value.
subject value
1.7672 bits per byteBits per byte of the subject.
comparator
Qwen2.5-1.5B continued with its own tokenizerModel whose value is comparator_value.
comparator value
1.5157 bits per byteBits per byte of the comparator.
evaluation text
math_clean statementsText scored.
language
EnglishLanguage of the text.
setting
after 500,170,752 Qwen tokens of the same document sequence, constant learning rate before any decay anneal, one run per armTraining both models received.

Experimental provenance

Method and evaluation protocol
Next-token negative log-likelihood in bits summed over non-overlapping 2048-token windows of each document, the first token of each window unscored, divided by the documents' UTF-8 bytes; bf16 on Apple silicon, 128 GB unified memory; arm-B tokenizers loaded through a whitespace canary and the frozen APT4 reference.
Dataset
The first 200 of 800 math_clean statements; 800 for the untouched model and arm A at 100M.Version: unspecified · Access: public
Reported results
Arm B 1.7671920928889029, arm A 1.5157282441289601; gap (arm B minus arm A) +0.2515; continued-pretraining effect on the same text (arm A minus untouched) +0.0135.
Uncertainty and replication
One scoring pass per model; no interval.
Limitations
One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept. The math_clean values of the untouched model and of arm A at 100M are over 800 statements and all others over the first 200; the rows were not re-scored.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

math_clean statements

Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.

Key math_clean_statements · version 1c1e659f-9adf-49a7-8976-649eb293aaa7

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

related

1832c8a3-4857-4668-b703-73fdb0289e46

An English capability statement about the transplanted models; this residual is measured at small scale after matched continued pretraining.