Finding

Sign in with GitHub
← Publications

Finding · P133 · Author-curated

After 0.5B tokens of the same continued pretraining, digit-probe accuracy is 0.4756 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.5167 for Qwen2.5-1.5B with its own tokenizer, a paired difference of -0.0411 (95% interval -0.0645 to -0.0177, sign-test p 0.0008).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired accuracy difference

metric
Digit-probe accuracyQuantity compared.
subject
Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continuedModel whose accuracy is subject_value.
subject value
0.4756 accuracyAccuracy of the subject.
comparator
Qwen2.5-1.5B continued with its own tokenizerModel whose accuracy is comparator_value.
comparator value
0.5167 accuracyAccuracy of the comparator.
difference
-0.0411 accuracySubject accuracy minus comparator accuracy over the same pairs.
interval low
-0.0645 accuracyLower bound of the 95% interval.
interval high
-0.0177 accuracyUpper bound of the 95% interval.
sign test p
0.0008 probabilityExact two-sided sign-test p-value over discordant pairs.
pairs
900 pairsItem-format pairs compared.
setting
after 500,170,752 Qwen tokens of the same document sequence, constant learning rate before any decay anneal, one run per armTraining both models received.

Experimental provenance

Method and evaluation protocol
300 synthetic arithmetic items from E2 in three prompt formats (bare English, bare Polish, Polish with spaced digits), 4-shot plain-text prompts, greedy decoding with batch 24, completion cut at the first newline, graded by E2's dual-locale grader after its 31-case golden suite; accuracy pooled over the 900 item-format pairs. Paired comparison: difference of accuracies over the same pairs, normal-approximation 95% interval from the discordant counts, exact two-sided sign test.
Dataset
300 items of E2's synthetic arithmetic probes (120 addition, 90 subtraction, 50 comparison, 40 sorting) in three prompt formats, frozen in configs/probe_digits_spec.json.Version: unspecified · Access: public
Reported results
Pairs 900; correct only for arm A 77, only for arm B 40; difference -0.04111111111111111; interval [-0.06451332152497036, -0.01770890069725187]; p 0.0007995844433963207.
Uncertainty and replication
The interval and p-value reflect item sampling in one run per arm, not run-to-run variance; pairs share items across three formats.
Evidence references
armA/ckpt_00500M.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/9a57aa40f4e774136f51b6730c33266d59094a45/experiments/E05-control-arm/results/probe_digits_raw/armA/ckpt_00500M.jsonlarmB/ckpt_00500M.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/9a57aa40f4e774136f51b6730c33266d59094a45/experiments/E05-control-arm/results/probe_digits_raw/armB/ckpt_00500M.jsonlprobe_digits_armA.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/9a57aa40f4e774136f51b6730c33266d59094a45/experiments/E05-control-arm/results/probe_digits_armA.jsonlprobe_digits_armB.jsonl · publichttps://github.com/stw2/tokenizer-science-tax/blob/9a57aa40f4e774136f51b6730c33266d59094a45/experiments/E05-control-arm/results/probe_digits_armB.jsonlanalysis.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/9a57aa40f4e774136f51b6730c33266d59094a45/experiments/E05-control-arm/results/analysis.jsonmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/9a57aa40f4e774136f51b6730c33266d59094a45/experiments/E05-control-arm/results/metrics.json
Limitations
One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired accuracy difference

Over the same item-format pairs, the subject's accuracy minus the comparator's accuracy is the difference, with a 95% interval from the discordant pairs and an exact two-sided sign-test p-value.

Key paired_accuracy_difference · version 683a738d-67cc-4630-b105-e5ad71c2d77a

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Exact references

related

36da68d9-97cf-4575-ad42-ed56e0e895e9

The GSM8K difference between the 11B transplanted model and its counterpart, of the same order.