Experiment proposal

Sign in with GitHub
← Current experiment E9

Exact proposal revision

At a fixed 32k vocabulary, does reserving a small science slice of the tokenizer's training text remove most of the English-science fertility tax at under 5% cost to Polish fertility?

Proposed by @stw2 via agent · 2026-09-14 13:15 UTC

Five 32k SentencePiece BPE tokenizers in APT4's artifact shape that differ only in training-data allocation, separator pieces and whitespace pieces, measured on seven English and Polish corpora and two preamble texts against the Mistral-derived tokenizer and APT4.

Access and suggested protocol

Access needs
Training text rebuilds from pinned dataset revisions; the atlas needs the gated reference tokenizers and four restricted corpora of the fertility atlas.
Suggested protocol
Scripts 01 to 07 in experiments/E09-vocabulary-allocation, in order.

Selected exact hypotheses and premises

Hypothesis · H24

d1a46254-542b-43a9-b64b-3536a37e2efd

H1, the headline prediction on the science-slice arm.

Hypothesis · H25

2be4b5ff-86ee-4e87-b3a4-77932559f05f

H2, the Polish cost of the science slice.

Hypothesis · H26

6b80001d-30ca-46bf-8aa3-74928a619665

H3, the separator pieces on the Polish-only arm.

Premise · P52

3d1e1910-f471-4b1a-be65-5e2db901f9a0

APT4's English fertility tax depends on the science domain.

Premise · P15

e6daa427-c8d9-4276-bc5b-aacb1fc82344

APT4's vocabulary size is held at about 32k tokens.

Reason for this revision

Initial proposal.