Finding

Sign in with GitHub
← Publications

Finding · P131 · Author-curated

Before any continued pretraining, digit-probe accuracy is 0.0133 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.5867 for untouched Qwen2.5-1.5B.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Digit-probe accuracyQuantity compared.
subject value
0.0133 accuracyAccuracy of the subject.
comparator
Qwen2.5-1.5BModel whose value is comparator_value.
comparator value
0.5867 accuracyAccuracy of the comparator.
scope
900 pairsItem-format pairs scored.

Experimental provenance

Method and evaluation protocol
300 synthetic arithmetic items from E2 in three prompt formats (bare English, bare Polish, Polish with spaced digits), 4-shot plain-text prompts, greedy decoding with batch 24, completion cut at the first newline, graded by E2's dual-locale grader after its 31-case golden suite; accuracy pooled over the 900 item-format pairs.
Dataset
300 items of E2's synthetic arithmetic probes (120 addition, 90 subtraction, 50 comparison, 40 sorting) in three prompt formats, frozen in configs/probe_digits_spec.json.Version: unspecified · Access: public
Reported results
Transplant at time zero 0.0133 (Wilson 95% [0.0076, 0.0232]); untouched 0.5867 (Wilson 95% [0.5542, 0.6184]).
Uncertainty and replication
Wilson 95% intervals over the 900 pairs, which are not independent across formats.
Limitations
Single model each. The time-zero row used the tokenizer loader of 6 July 2026; its prompt dump is identical to every later arm-B dump.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Accuracies are overall.acc of results/probe_digits_armA.jsonl and probe_digits_armB.jsonl; analysis.json pools the rounded per-format accuracies and gives 0.5866 for the untouched model.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, before training

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, before any further training.

Key apt4_fvt_time_zero · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Digit-probe accuracy

Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.

Key digit_probe_accuracy · version fbf37827-4b31-402b-8eba-0bb9d16d62c9

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Exact references

related

85dfefb0-eb84-4deb-b93c-7b494a10b283

Digit handling of a tokenizer can influence generation quality.