Bielik 1.5B v3
The 1.5B-parameter Bielik v3 model.
Key bielik_1_5b_v3 · version f5ada7c4-30c8-405f-8741-cd6fcfb58530
Cited claim
Sign in with GitHubCited claim · P24 · Author-curated
Relation: Supported by evidence
Our choice of FOCUS is supported by prior experimental results on earlier Bielik v3 models Ociepa et al. [2025b], where multiple embedding initialization strategies were systematically evaluated.
Among these approaches, FOCUS consistently demonstrated the best empirical performance. In particular, experiments conducted on the Bielik 1.5B v3 model showed the lowest training loss after 4B tokens of continued pretraining, as well as leading results on the Open Polish LLM Leaderboard Wróbel et al. [2024], Ociepa et al. [2025c].
The evidence is reported in Ociepa et al. [2025b]; this paper gives no measurement of it.
Reuse the defining version and key when the meaning fits your assertion.
The 1.5B-parameter Bielik v3 model.
Key bielik_1_5b_v3 · version f5ada7c4-30c8-405f-8741-cd6fcfb58530
The choice is supported by the evidence stated in the evidence and result roles.
Key supported_by_evidence · version f5ada7c4-30c8-405f-8741-cd6fcfb58530
Fast Overlapping Token Combinations Using Sparsemax: an embedding initialisation for a replaced tokenizer's vocabulary.
Key focus_init · version 5a078e71-84ff-4039-9e32-7999a5f679f5
Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.