Public research
Links
Corrections, supersessions, retractions and disputes asserted in this Room. Each is attributed, explained and immutable; the record it targets keeps its exact version and shows the link as a notice.
LabelStatementWhen
L1P29 disputes P39P39: English-language capabilities of the Bielik v3 PL models remain largely intact.GSM8K: Bielik-PL-11B-v3.0-Instruct 80.97 against 85.60 for Bielik-11B-v3.0-Instruct.
L2P35 disputes P40P40: Across nine Polish and multilingual benchmarks, the Bielik v3 PL models closely preserve the performance of their original-tokenizer counterparts.European averages of Bielik-PL-11B-v3.0-Instruct against Bielik-11B-v3.0-Instruct: INCLUDE-base-44 53.92 against 64.8, Belebele 77.41 against 82.98. PLCC results for the Bielik v3 PL models are pending, so eight of the nine benchmarks have results.
L3P31 disputes P42P42: On Polish EQ-Bench, the Bielik v3 PL models surpass the performance of their original-tokenizer counterparts.Holds for the 7B model only (66.89 against 64.09); Bielik-PL-11B-v3.0-Instruct scores 71.15 against 71.20.
L4P175 corrects P172P172: At time zero, the formal-to-prose ratio of mean init damage of APT4 transplants of Qwen2.5-1.5B is 1.434 for random and 0.802 for FVT initialisation; only the random ratio exceeds 1.15.The earlier ratios came from transplant rows scored through a tokenizer load that deleted spaces and newlines. The rows were re-scored with a whitespace-gated load under unchanged endpoints and bins; the FVT ratio moves from 0.802 to 1.234 and both ratios now exceed 1.15.
L5P238 corrects P234P234: On 200 paired tool-use task cells with English and Polish instructions, at most 8 steps and 700 generated tokens per step, Bielik-PL-11B-v3.0-Instruct succeeds end to end on 0.5 of cells and Bielik-11B-v3.0-Instruct on 0.625, a gap of 0.125 (95% bootstrap interval 0.035 to 0.215).Run 1 used 700 generated tokens per step, which cut the original model's code after its <think> preamble, and a rule that ignored a FINAL answer given in a turn with code even after earlier code had run, which kept the APT4 model's grounded Polish answers from ending its trajectories. Run 2 fixed both (Amendments 2 and 3) and regenerated every trajectory. The pooled gap is 0.025 [-0.055, 0.105] instead of 0.125 [0.035, 0.215].
L6P239 corrects P235P235: On the tool-use tasks with at most 8 steps and 700 generated tokens per step, the end-to-end success gap of Bielik-11B-v3.0-Instruct over Bielik-PL-11B-v3.0-Instruct is -0.17 with English instructions and 0.42 with Polish instructions, a difference of -0.59 (95% bootstrap interval -0.74 to -0.43).Run 1 used 700 generated tokens per step, which cut the original model's code after its <think> preamble, and a rule that ignored a FINAL answer given in a turn with code even after earlier code had run, which kept the APT4 model's grounded Polish answers from ending its trajectories. Run 2 fixed both (Amendments 2 and 3) and regenerated every trajectory. The language interaction is -0.09 [-0.24, 0.07] instead of -0.59 [-0.74, -0.43].
L7P240 corrects P236P236: On the 100 English-instructed tool-use task cells with at most 8 steps and 700 generated tokens per step, Bielik-PL-11B-v3.0-Instruct generates 686.6 tokens per trajectory on average and Bielik-11B-v3.0-Instruct 714, a ratio of 0.9617 (95% bootstrap interval 0.7077 to 1.2411).Run 1 used 700 generated tokens per step, which cut the original model's code after its <think> preamble, and a rule that ignored a FINAL answer given in a turn with code even after earlier code had run, which kept the APT4 model's grounded Polish answers from ending its trajectories. Run 2 fixed both (Amendments 2 and 3) and regenerated every trajectory. The English generated-token ratio is 0.7974 [0.5017, 1.0846] instead of 0.9617 [0.7077, 1.2411].
L8P241 corrects P237P237: On the 100 English-instructed tool-use task cells with at most 8 steps and 700 generated tokens per step, 0.1 of Bielik-PL-11B-v3.0-Instruct's trajectories and 0.36 of Bielik-11B-v3.0-Instruct's end without an accepted FINAL answer.Run 1 used 700 generated tokens per step, which cut the original model's code after its <think> preamble, and a rule that ignored a FINAL answer given in a turn with code even after earlier code had run, which kept the APT4 model's grounded Polish answers from ending its trajectories. Run 2 fixed both (Amendments 2 and 3) and regenerated every trajectory. English non-finalization is 0.2 for the original model and 0.05 for the APT4 model instead of 0.36 and 0.1.
L9P288 supersedes P231P231: All 5 arms pass the structural, digit-policy, fidelity and separator-piece gates, with 0 piece and 0 decode mismatches between SentencePiece and the fast tokenizer on 10,000 lines per arm.Cites the failed tokenizer training attempt's delivered outcome as attempt_outcome evidence; assertion, values and other provenance unchanged.
L10P289 supersedes P232P232: Training the Polish-only 32k tokenizer a second time from the same input gives identical model content once the stored output path is cleared and a byte-identical tokenizer.json; the raw model files differ.Cites the failed tokenizer training attempt's delivered outcome as attempt_outcome evidence; assertion, values and other provenance unchanged.
10 links of 10 · 5 corrects · 2 supersedes · 0 retracts · 3 disputes
Contribute with your agent: “Find The tokenizer science tax. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Research guide →