Finding

Sign in with GitHub
← Publications

Finding · P135 · Author-curated

On PES examination questions, likelihood accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer is 0.2280, 0.2320, 0.2320, 0.2200, 0.2280 and 0.2560, and of Qwen2.5-1.5B with its own tokenizer 0.2480, 0.2720, 0.2880, 0.2720, 0.2640 and 0.2560, after 0, 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Arm trajectories

metric
Multiple-choice likelihood accuracyQuantity measured.
benchmark
PES examination questionsItems scored.
subject
Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continuedArm whose values are subject_values.
subject values
0.2280; 0.2320; 0.2320; 0.2200; 0.2280; 0.2560Accuracy of the subject arm at each milestone.
comparator
Qwen2.5-1.5B continued with its own tokenizerArm whose values are comparator_values.
comparator values
0.2480; 0.2720; 0.2880; 0.2720; 0.2640; 0.2560Accuracy of the comparator arm at each milestone.
milestones
0; 100,139,008; 200,278,016; 300,417,024; 400,031,744; 500,170,752 tokensTraining tokens of each value.
items
250 itemsItems scored per model.
language
PolishLanguage of the items.

Experimental provenance

Method and evaluation protocol
For each item, the summed log-probability of the answer letter after the question and 'Odpowiedź:'; the highest-scoring letter is the prediction; first items of each set in file order.
Dataset
The first 250 items of E3's PES examination questions.Version: unspecified · Access: restricted
Reported results
Arm B 0.228, 0.232, 0.232, 0.22, 0.228, 0.256; arm A 0.248, 0.272, 0.288, 0.272, 0.264, 0.256.
Uncertainty and replication
One scoring pass per model; no interval; accuracies lie near the chance level of four-option items.
Limitations
Polish examination items only, so the probe cannot show an English collapse. One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept. The time-zero arm-B rows were scored with an earlier tokenizer loader that rebuilt from the checkpoint's own tokenizer.json and were not re-scored; the encodings are probably, not verifiably, equal.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Arm trajectories

subject_values and comparator_values give the metric of the subject and comparator arms at the listed training-token milestones; at 0 tokens each arm is its starting model.

Key arm_trajectories · version 4cd12925-1344-4a79-9c5b-4c2c52b50ce3

PES examination questions

250 multiple-choice questions of the Polish State Specialization Examination 2018 to 2022, allocated proportionally over 57 medical specializations.

Key pes_questions · version a164a223-a32f-4fdf-a694-c1d5cda9b77d

Qwen2.5-1.5B continued with its own tokenizer

Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.

Key original_tokenizer_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584

Multiple-choice likelihood accuracy

Share of multiple-choice items on which the answer letter with the highest summed log-probability after the question is the gold letter.

Key mcq_accuracy · version 4cd12925-1344-4a79-9c5b-4c2c52b50ce3

Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer, continued

Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.

Key apt4_fvt_cpt_arm · version 9a53082a-65e8-4c6a-82be-33dda5810584