Finding

Sign in with GitHub
← Publications

Finding · P114 · Author-curated

Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens lowers bits per byte on Polish Wikipedia science articles from 1.0930 to 0.8782.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Bits per byteQuantity compared.
subject
Qwen2.5-1.5B continued with its own tokenizerModel whose value is subject_value.
subject value
0.8782 bits per byteBits per byte of the subject.
comparator
Qwen2.5-1.5BModel whose value is comparator_value.
comparator value
1.0930 bits per byteBits per byte of the comparator.
evaluation text
Polish Wikipedia science articlesText scored.
language
PolishLanguage of the text.
setting
500,170,752 Qwen tokens of continued pretraining at a constant learning rate before any decay anneal, one runTraining of the subject.
scope
first 200 documents of the corpusDocuments scored.

Experimental provenance

Method and evaluation protocol
Next-token negative log-likelihood in bits summed over non-overlapping 2048-token windows of each document, the first token of each window unscored, divided by the documents' UTF-8 bytes; bf16 on Apple silicon, 128 GB unified memory; arm-B tokenizers loaded through a whitespace canary and the frozen APT4 reference.
Dataset
The first 200 documents of E1's Polish Wikipedia science corpus.Version: unspecified · Access: public
Reported results
Arm A 0.8781854339415027, untouched 1.0929525011709564; continued-pretraining effect (arm A minus untouched) -0.2148.
Uncertainty and replication
One scoring pass per model; no interval. The untouched model's row was scored twice with identical values.
Limitations
One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Polish Wikipedia science articles

Opening paragraphs, up to 900 words, of 229 Polish Wikipedia articles selected by science keywords.

Key polish_wikipedia_science · version 8e524448-ccc5-405f-830d-82c821643387

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

related

d664a095-2689-445e-ac40-a85c6202a460

The continued pretraining whose effect this isolates at small scale.