Finding

Sign in with GitHub
← Publications

Finding · P103 · Author-curated

Fast Vocabulary Transfer of APT4 into Qwen2.5-1.5B initialises 31,741 pieces as means of their Qwen re-encodings, which average 2.4116 Qwen pieces; 256 byte-fallback pieces and 3 special pieces receive fixed rows, and 0 pieces have an empty re-encoding.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Initialisation counts

method
Frequency-based Vocabulary TransferInitialisation method.
source model
Qwen2.5-1.5BModel whose embeddings are the source.
target tokenizer
APT4 tokenizerTokenizer whose pieces are initialised.
method pieces
31741 piecesPieces initialised from their re-encodings.
mean source pieces
2.4116 piecesMean number of source pieces per re-encoded piece.
byte pieces
256 piecesByte-fallback pieces given the mean embedding.
special pieces
3 piecesSpecial pieces given fixed rows.
empty pieces
0 piecesPieces whose re-encoding is empty.

Experimental provenance

Method and evaluation protocol
Each APT4 piece, with the word-boundary marker replaced by a space, is re-encoded with the Qwen tokenizer and its embedding set to the mean of those Qwen rows; <s> and </s> take the <|endoftext|> row; <unk> and byte pieces take the mean embedding; embeddings are tied.
Dataset
The APT4 vocabulary of Bielik-PL-11B-v3.0-Instruct and the Qwen2.5-1.5B embedding matrix.Version: unspecified · Access: restricted
Reported results
fvt 31741, byte 256, special 3, empty 0, mean Qwen pieces per re-encoded piece 2.411581235625847.
Uncertainty and replication
Exact counts.
Limitations
The statistics file was not committed in the archive; it was copied from the scoring machine's time-zero model directory. The script as archived did not write the eos fix of the model that was trained.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Initialisation counts

Initialising the target tokenizer's embeddings in the source model by the method assigns method_pieces pieces from their re-encodings, averaging mean_source_pieces source pieces each, gives byte_pieces byte-fallback and special_pieces special pieces fixed rows, and leaves empty_pieces pieces without a re-encoding.

Key initialisation_counts · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Frequency-based Vocabulary Transfer

An embedding initialisation for a replaced tokenizer's vocabulary.

Key fvt_init · version 01087ca0-917c-46f8-a25e-ebc786018d72

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Exact references

related

01087ca0-917c-46f8-a25e-ebc786018d72

The initialisation method, applied here with plain means of the constituent rows.

related

5a078e71-84ff-4039-9e32-7999a5f679f5

The initialisation the Bielik v3 PL models use instead.