Vocabulary size
Number of tokens in a tokenizer's vocabulary.
Key vocabulary_size · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e
Cited claim
Sign in with GitHubCited claim · P14 · Author-curated
Relation: Metric comparison
The original Bielik v3 tokenizer employed a vocabulary of 32,128 tokens.
To address this issue, the Bielik v3 PL models adopt a dedicated Polish tokenizer with a comparable vocabulary size of 32,000 tokens.
Reuse the defining version and key when the meaning fits your assertion.
Number of tokens in a tokenizer's vocabulary.
Key vocabulary_size · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e
The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.
Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e
Tokenizer of the original Bielik v3 models, derived from Mistral's.
Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7
Polish-optimised tokenizer of the Bielik v3 PL models.
Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c
Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.