Cited claim

Sign in with GitHub
← Publications

Cited claim · P25 · Author-curated

Vocabulary adaptation of the Bielik v3 PL models uses a 20B-token subset sampled from the original Bielik 11B v3 corpus.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Trained on

subject
Bielik v3 PL modelsModels trained.
stage
Vocabulary adaptationTraining stage.
token budget
20000000000 tokensTokens used.
data source
original Bielik 11B v3 corpusCorpus the data is sampled from.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Section 4.1, p. 4

    The training data consists of a 20B-token subset sampled from the original Bielik 11B v3 corpus

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Trained on

The subject's training stage uses the data source and token budget stated.

Key trained_on · version d664a095-2689-445e-ac40-a85c6202a460

Vocabulary adaptation

Adaptation of a pretrained model to a replaced tokenizer's vocabulary.

Key vocabulary_adaptation · version 95d5752d-3660-4d7c-97d5-f23b933cf433

Bielik v3 PL models

The 11B and 7B Bielik v3 models with the APT4 tokenizer.

Key bielik_v3_pl_models · version 3da77265-eb29-46f1-90e2-ccc03ac36917