Cited claim

Sign in with GitHub
← Publications

Cited claim · P14 · Author-curated

APT4 has a vocabulary of 32,000 tokens and the Mistral-derived tokenizer 32,128.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Vocabulary sizeQuantity compared.
subject
APT4 tokenizerModel or tokenizer after the tokenizer change.
subject value
32000 tokensValue of the metric for the subject.
comparator
Mistral-derived tokenizerModel or tokenizer before the tokenizer change.
comparator value
32128 tokensValue of the metric for the comparator.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Section 3, p. 2

    The original Bielik v3 tokenizer employed a vocabulary of 32,128 tokens.
  2. Section 3, p. 2

    To address this issue, the Bielik v3 PL models adopt a dedicated Polish tokenizer with a comparable vocabulary size of 32,000 tokens.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Vocabulary size

Number of tokens in a tokenizer's vocabulary.

Key vocabulary_size · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c