Finding

Sign in with GitHub
← Publications

Finding · P174 · Author-curated

31,820 of the 32,000 APT4 pieces are in-vocabulary words of fastText auxiliary embeddings trained on 48,050,703 APT4 tokens of Polish FineWeb2-HQ text.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Vocabulary coverage

tokenizer
APT4 tokenizerTokenizer whose pieces are counted.
auxiliary model
gensim fastText, skip-gram, 200 dimensions, character n-grams 3 to 6, 3 epochsEmbedding model the pieces are looked up in.
training text
the first 200 MB of Polish FineWeb2-HQ training text, tokenized with APT4, one piece per wordText the auxiliary model was trained on.
training tokens
48050703 tokensTokens in the training text.
count in vocabulary
31820 piecesPieces that are words of the auxiliary model.
count total
32000 piecesPieces of the tokenizer.

Experimental provenance

Method and evaluation protocol
Count of APT4 pieces present in the trained fastText vocabulary; the other pieces get vectors synthesised from character n-grams.
Dataset
The first 209,715,204 bytes of Polish FineWeb2-HQ training text.Version: unspecified · Access: restricted
Reported results
31820 of 32,000 pieces in vocabulary; FOCUS fell back to FVT for 0 new tokens.
Uncertainty and replication
Exact count for one training run; the training is not bit-reproducible with 4 worker threads, so the matrix is frozen by digest.
Limitations
Monolingual Polish auxiliary space. The locked design specified 300 dimensions, 5 epochs and one worker; the run used 200 dimensions, 3 epochs and 4 workers. The auxiliary fastText text (the first 209,715,204 bytes of the Polish FineWeb2-HQ training text) and the Polish FineWeb2-HQ holdout were drawn from the same source without a shared exclusion: a line check in the port found that 3 of the 180 scored holdout documents with lines of 200 to 4,000 bytes share at least one such line with the auxiliary text, and none shares half of them.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Values from data/aux/aux_manifest.json, aux_rows_in_ft_vocab and corpus_tokens.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Vocabulary coverage

count_in_vocabulary of count_total pieces of the tokenizer are words of the auxiliary_model trained on training_tokens tokens of the training_text.

Key vocabulary_coverage · version daa6c044-3484-44d0-98d8-e07276dbbd8d

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Exact references

related

5a078e71-84ff-4039-9e32-7999a5f679f5

The auxiliary embedding space the method needs and the paper does not specify.