Experiment

Sign in with GitHub
← Experiments

Experiment · E21

Does the science-slice tokenizers' lower English-science fertility tax carry over to lower bits per byte on formal texts after transplant and continued pretraining, and at what Polish cost?

Available · Proposed by @stw2 · Unassigned

Five arms of Qwen2.5-0.5B, each continued-pretrained for 1B tokens on the control-arm experiment's deterministic document sequence: Fast Vocabulary Transfer transplants of the Polish-only 32k tokenizer, the Polish-only 32k tokenizer with separator pieces, the science-slice 32k tokenizer and the science-slice 32k tokenizer with whitespace pieces, and the model with its own tokenizer as the reference. Primary endpoint: the residual ratio of the whitespace science-slice arm to the Polish-only arm, pooled over formal texts. Secondary: that ratio against the same tokenizers' fertility-tax ratio, the pooled Polish cost, and matched-step and matched-byte comparisons with byte-anchored windows.

Prerequisites and protocol

Access needs
Training needs CUDA GPUs with 24 GB, one per arm. The packed training streams are rebuilt by script from pinned dataset revisions. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script.
Suggested protocol
Commit a design before training, with the bin: a residual ratio of at most 0.5 confirms. Reuse scripts/02_pack.py, 03_transplant_fvt.py and 04_train.py of experiments/E05-control-arm and scripts/00_bootstrap_cuda.sh of experiments/E06-math-code-rescue; the arm tokenizers ship in experiments/E09-vocabulary-allocation. Feared artifact: tokenizer-training-corpus leakage. The science arms' vocabularies were built on English science text, and the Polish training text of every arm tokenizer overlaps the Polish evaluation holdout; audit n-gram overlap between each tokenizer's training text and every evaluation text, and add a held-out formal domain that no tokenizer saw (Lean or Coq proof scripts, or SMILES strings). Kill criteria: a ratio of at least 0.9 means the allocation buys nothing in loss; at most 0.2 means full transfer.

Accepted plan

No accepted plan. Available work need not have a complete protocol or source commit.

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

No attempt registered. Work status and findings are independent of attempts.

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →