Exact proposal revision
At a fixed 32k vocabulary, does reserving a small science slice of the tokenizer's training text remove most of the English-science fertility tax at under 5% cost to Polish fertility?
Five 32k SentencePiece BPE tokenizers in APT4's artifact shape that differ only in training-data allocation, separator pieces and whitespace pieces, measured on seven English and Polish corpora and two preamble texts against the Mistral-derived tokenizer and APT4.
Access and suggested protocol
- Access needs
- Training text rebuilds from pinned dataset revisions; the atlas needs the gated reference tokenizers and four restricted corpora of the fertility atlas.
- Suggested protocol
- Scripts 01 to 07 in experiments/E09-vocabulary-allocation, in order.
Selected exact hypotheses and premises
Hypothesis · H24
d1a46254-542b-43a9-b64b-3536a37e2efdH1, the headline prediction on the science-slice arm.
Hypothesis · H26
6b80001d-30ca-46bf-8aa3-74928a619665H3, the separator pieces on the Polish-only arm.
Premise · P52
3d1e1910-f471-4b1a-be65-5e2db901f9a0APT4's English fertility tax depends on the science domain.
Premise · P15
e6daa427-c8d9-4276-bc5b-aacb1fc82344APT4's vocabulary size is held at about 32k tokens.
Reason for this revision
Initial proposal.