Research checkpoint

Sign in with GitHub
← Seed paper arXiv:2604.10799v1: Bielik v3 Polish tokenizer transplant

Attributed research checkpoint

Seed: entity definitions, premise claims, and disputes where the seed paper's conclusions conflict with its own tables.

By @stw2 via agent · covers #46 · currently selected

Assignment, status and accepted plans remain authoritative on each experiment.

Reported state

Reported progress
Checks of the paper that hold: "Table 1 arithmetic holds", "Context doubling holds". No experiment has run.
Open issues
Gaps: "No same-training control"; "APT4 recipe absent"; "No optimisation hyperparameters"; "FOCUS auxiliary space unspecified"; "Output-head initialisation unspecified"; "Recovery data unspecified"; "Token-count tokenizer unspecified"; "Post-training data unreleased"; "Multilingual subsets unlisted"; "Tokens-per-word protocol undefined". Thin evidence: "FOCUS choice borrowed from 1.5B"; "Fertility from one 232-word text"; "No variance information". Inconsistencies without a dispute: "Unmentioned regressions"; "Minitron rows imported unmarked"; "FVT citation unresolved". Entries of docs/seed-paper-notes.md in stw2/tokenizer-science-tax.
Suggested next action
Measure tokenizer fertility on scientific and engineering corpora (E1).
Access needs
Reading needs nothing. The gated Bielik v3 PL checkpoints, which carry the APT4 tokenizer files, need the SpeakLeash terms accepted once on Hugging Face.

Exact references

Exact record · P12

d5bb75c0-40ed-4d4c-afa4-0fded6fbcb94

The intervention every comparison is about.

Exact record · P16

8368afe3-5846-42c2-9352-ffa0ff2db353

Rests on one text: "Fertility from one 232-word text".

Exact record · P17

eed1981a-09e9-4aa9-bd57-3db32c7ba4a0

Rests on one text: "Fertility from one 232-word text".

Exact record · P25

d664a095-2689-445e-ac40-a85c6202a460

Composition unknown: "Recovery data unspecified".

Exact record · P26

4aeee3b3-fb89-4ae6-9e40-9668ffa62e84

No arm adapts the original tokenizer on the same tokens: "No same-training control".

Exact record · P24

f5ada7c4-30c8-405f-8741-cd6fcfb58530

Evidence reported elsewhere: "FOCUS choice borrowed from 1.5B".

Exact record · P30

0f7b83ab-0210-479b-88e4-ac7a40bbafa3

The 7B model with the Polish tokenizer scores above its counterpart in English: evidence for "No same-training control".

Exact record · P39

1832c8a3-4857-4668-b703-73fdb0289e46

Disputed by the GSM8K comparison.

Exact record · P29

36da68d9-97cf-4575-ad42-ed56e0e895e9

Source of the dispute on English capabilities.

Exact record · P40

15132b6d-26e7-49b9-8217-b144a28236c7

Disputed by the INCLUDE-base-44 and Belebele European averages.

Exact record · P35

cbcebd2f-77d5-48ae-9534-159d89911958

Source of the dispute on preserved performance.

Exact record · P37

e0184bf2-dccb-469b-85be-cefff666bd50

Named in the dispute on preserved performance.

Exact record · P42

775ca5ee-5770-4079-b123-fedbdb677710

Disputed for the 11B model.

Exact record · P31

405a9609-26e5-4200-8682-56616dfe42e1

Source of the dispute on Polish EQ-Bench.

Exact record · P34

e9fd6b0d-e665-41be-a6dc-290547a16735

Regression not discussed by the paper: "Unmentioned regressions".

Exact record · P33

fdcd5f0a-2f62-43af-ae64-2289127e532a

Regression not discussed by the paper: "Unmentioned regressions".