Finding

Sign in with GitHub
← Publications

Finding · P105 · Author-curated

After 0.5B tokens of the same continued pretraining, bits per byte on Polish Wikipedia science articles is 0.9116 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.8782 for Qwen2.5-1.5B with its own tokenizer.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Bits per byteQuantity compared.
subject
Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continuedModel whose value is subject_value.
subject value
0.9116 bits per byteBits per byte of the subject.
comparator
Qwen2.5-1.5B continued with its own tokenizerModel whose value is comparator_value.
comparator value
0.8782 bits per byteBits per byte of the comparator.
evaluation text
Polish Wikipedia science articlesText scored.
language
PolishLanguage of the text.
setting
after 500,170,752 Qwen tokens of the same document sequence, constant learning rate before any decay anneal, one run per armTraining both models received.
scope
first 200 documents of the corpusDocuments scored.

Experimental provenance

Method and evaluation protocol
Next-token negative log-likelihood in bits summed over non-overlapping 2048-token windows of each document, the first token of each window unscored, divided by the documents' UTF-8 bytes; bf16 on Apple silicon, 128 GB unified memory; arm-B tokenizers loaded through a whitespace canary and the frozen APT4 reference.
Dataset
The first 200 documents of E1's Polish Wikipedia science corpus.Version: unspecified · Access: public
Reported results
Arm B 0.9116095430921759, arm A 0.8781854339415027; gap (arm B minus arm A) +0.0334; continued-pretraining effect on the same text (arm A minus untouched) -0.2148.
Uncertainty and replication
One scoring pass per model; no interval.
Limitations
One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Polish Wikipedia science articles

Opening paragraphs, up to 900 words, of 229 Polish Wikipedia articles selected by science keywords.

Key polish_wikipedia_science · version 8e524448-ccc5-405f-830d-82c821643387

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

related

d664a095-2689-445e-ac40-a85c6202a460

The transplanted models' continued pretraining, which both arms here receive in matched form.