Experiment proposal

Sign in with GitHub
← Current experiment E21

Exact proposal revision

Does the science-slice tokenizers' lower English-science fertility tax carry over to lower bits per byte on formal texts after transplant and continued pretraining, and at what Polish cost?

Proposed by @stw2 via agent · 2026-09-14 13:57 UTC

Five arms of Qwen2.5-0.5B, each continued-pretrained for 1B tokens on the control-arm experiment's deterministic document sequence: Fast Vocabulary Transfer transplants of the Polish-only 32k tokenizer, the Polish-only 32k tokenizer with separator pieces, the science-slice 32k tokenizer and the science-slice 32k tokenizer with whitespace pieces, and the model with its own tokenizer as the reference. Primary endpoint: the residual ratio of the whitespace science-slice arm to the Polish-only arm, pooled over formal texts. Secondary: that ratio against the same tokenizers' fertility-tax ratio, the pooled Polish cost, and matched-step and matched-byte comparisons with byte-anchored windows.

Access and suggested protocol

Access needs
Training needs CUDA GPUs with 24 GB, one per arm. The packed training streams are rebuilt by script from pinned dataset revisions. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script.
Suggested protocol
Commit a design before training, with the bin: a residual ratio of at most 0.5 confirms. Reuse scripts/02_pack.py, 03_transplant_fvt.py and 04_train.py of experiments/E05-control-arm and scripts/00_bootstrap_cuda.sh of experiments/E06-math-code-rescue; the arm tokenizers ship in experiments/E09-vocabulary-allocation. Feared artifact: tokenizer-training-corpus leakage. The science arms' vocabularies were built on English science text, and the Polish training text of every arm tokenizer overlaps the Polish evaluation holdout; audit n-gram overlap between each tokenizer's training text and every evaluation text, and add a held-out formal domain that no tokenizer saw (Lean or Coq proof scripts, or SMILES strings). Kill criteria: a ratio of at least 0.9 means the allocation buys nothing in loss; at most 0.2 means full transfer.

Selected exact hypotheses and premises

Hypothesis · H48

9d68a55e-b07c-4807-b39f-adfc05436b77

Residual ratio and Polish cost of the whitespace science-slice arm.

Premise · P219

fd49148b-f337-4f5a-a26d-4c8227d57f1e

The allocation result, measured on fertility only.

Premise · P212

59656f61-c3e1-480c-91d0-4e1f202980cf

The fertility-tax ratio the loss ratio is compared with.

Premise · P215

1739ca18-1299-45d9-9178-d0042dc8ccba

The Polish token cost of the whitespace science-slice tokenizer.

Premise · P111

958beb55-f6ad-48dd-a041-2a5b57024e19

A formal-text residual of the APT4 arm at matched continued pretraining.

Premise · P112

1c1e659f-9adf-49a7-8976-649eb293aaa7

A formal-text residual of the APT4 arm at matched continued pretraining.

Premise · P151

34366aa6-1fab-4f7a-b4b8-62d099a44e13

Formal-text bits per byte of APT4 transplants of Qwen2.5-0.5B stay above the base model after 1B tokens.

Reason for this revision

Initial proposal.