Finding

Sign in with GitHub
← Publications

Finding · P132 · Author-curated

Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens lowers digit-probe accuracy from 0.5867 to 0.5167.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Digit-probe accuracyQuantity compared.
subject
Qwen2.5-1.5B continued with its own tokenizerModel whose value is subject_value.
subject value
0.5167 accuracyAccuracy of the subject.
comparator
Qwen2.5-1.5BModel whose value is comparator_value.
comparator value
0.5867 accuracyAccuracy of the comparator.
setting
500,170,752 Qwen tokens of continued pretraining without a mathematics or code slice, one runTraining of the subject.
scope
900 pairsItem-format pairs scored.

Experimental provenance

Method and evaluation protocol
300 synthetic arithmetic items from E2 in three prompt formats (bare English, bare Polish, Polish with spaced digits), 4-shot plain-text prompts, greedy decoding with batch 24, completion cut at the first newline, graded by E2's dual-locale grader after its 31-case golden suite; accuracy pooled over the 900 item-format pairs.
Dataset
300 items of E2's synthetic arithmetic probes (120 addition, 90 subtraction, 50 comparison, 40 sorting) in three prompt formats, frozen in configs/probe_digits_spec.json.Version: unspecified · Access: public
Reported results
Arm A at 500M 0.5167 (Wilson 95% [0.4841, 0.5492]); untouched 0.5867; arm A at 100M 0.6167, 300M 0.5489.
Uncertainty and replication
Wilson 95% intervals over the 900 pairs; the two models were not compared item by item.
Limitations
One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584