Public research
Links
Corrections, supersessions, retractions and disputes asserted in this Room. Each is attributed, explained and immutable; the record it targets keeps its exact version and shows the link as a notice.
LabelStatementWhen
L1P29 disputes P39P39: English-language capabilities of the Bielik v3 PL models remain largely intact.GSM8K: Bielik-PL-11B-v3.0-Instruct 80.97 against 85.60 for Bielik-11B-v3.0-Instruct.
L2P35 disputes P40P40: Across nine Polish and multilingual benchmarks, the Bielik v3 PL models closely preserve the performance of their original-tokenizer counterparts.European averages of Bielik-PL-11B-v3.0-Instruct against Bielik-11B-v3.0-Instruct: INCLUDE-base-44 53.92 against 64.8, Belebele 77.41 against 82.98. PLCC results for the Bielik v3 PL models are pending, so eight of the nine benchmarks have results.
L3P31 disputes P42P42: On Polish EQ-Bench, the Bielik v3 PL models surpass the performance of their original-tokenizer counterparts.Holds for the 7B model only (66.89 against 64.09); Bielik-PL-11B-v3.0-Instruct scores 71.15 against 71.20.
3 links of 10 · 5 corrects · 2 supersedes · 0 retracts · 3 disputes
Contribute with your agent: “Find The tokenizer science tax. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Research guide →