Experiment

Sign in with GitHub
← Experiments

Experiment · E9

At a fixed 32k vocabulary, does reserving a small science slice of the tokenizer's training text remove most of the English-science fertility tax at under 5% cost to Polish fertility?

Completed · Proposed by @stw2 · Assigned to @stw2

Five 32k SentencePiece BPE tokenizers in APT4's artifact shape that differ only in training-data allocation, separator pieces and whitespace pieces, measured on seven English and Polish corpora and two preamble texts against the Mistral-derived tokenizer and APT4.

Prerequisites and protocol

Access needs
Training text rebuilds from pinned dataset revisions; the atlas needs the gated reference tokenizers and four restricted corpora of the fertility atlas.
Suggested protocol
Scripts 01 to 07 in experiments/E09-vocabulary-allocation, in order.

Accepted plan

Plan accepted by @stw2

Five 32k SentencePiece BPE tokenizers differing in training-data allocation, audited for digit and separator handling and measured on a fertility atlas of seven corpora

retrospective

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts · 9

8 succeeded · 1 failed

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

  1. planned attempt · stw2/tokenizer-science-tax @ a77445db2440 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  2. planned attempt · stw2/tokenizer-science-tax @ a77445db2440 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  3. planned attempt · stw2/tokenizer-science-tax @ f52d5afd8e13 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  4. rerun attempt · stw2/tokenizer-science-tax @ 7a2cffbe78cf · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  5. planned attempt · stw2/tokenizer-science-tax @ 7a2cffbe78cf · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  6. planned attempt · stw2/tokenizer-science-tax @ 7a2cffbe78cf · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  7. planned attempt · stw2/tokenizer-science-tax @ 7a2cffbe78cf · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  8. planned attempt · stw2/tokenizer-science-tax @ 7a2cffbe78cf · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  9. planned attempt · stw2/tokenizer-science-tax @ 7a2cffbe78cf · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →