Execution attempt

Sign in with GitHub
← Experiment E9 · At a fixed 32k vocabulary, does reserving a small science slice of the tokenizer's training text remove most of the English-science fertility tax at under 5% cost to Polish fertility?

Execution attempt · A85 · planned

Trains the five 32k arms and converts each to a fast tokenizer; gates G1, G4 and G5 pass for every arm. Then retrains the Polish-only arm for the determinism gate G3, which fails: the retrained model file is not byte-identical, tokenizer.json is. The arm inputs of the two Polish-only arms are byte-identical to pl.txt and are not listed. The G3 comparison code of this invocation predates Amendment 2 and was not kept; the pinned commit holds the amended code.

Failed · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ f52d5afd8e13cbace7a438d4fd653deeb64de610

Reference checked 2026-09-14 13:15 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/02_train_tokenizers.py --g3
Working directory
experiments/E09-vocabulary-allocation
Configuration paths
None
Parameters
vocabulary 32,000; BPE; character coverage 0.9995; identity normalisation; byte fallback; split digits; maximum piece length 16; no input shuffling; separator pieces as user-defined symbols in four arms; whitespace-only pieces in one arm; num_threads 8
Environment
3.12.13; SentencePiece BPE trainer and Hugging Face tokenizers on CPU; no model weights; Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Delivered events

  1. Registered

    #1

    Trains the five 32k arms and converts each to a fast tokenizer; gates G1, G4 and G5 pass for every arm. Then retrains the Polish-only arm for the determinism gate G3, which fails: the retrained model file is not byte-identical, tokenizer.json is. The arm inputs of the two Polish-only arms are byte-identical to pl.txt and are not listed. The G3 comparison code of this invocation predates Amendment 2 and was not kept; the pinned commit holds the amended code.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Failed

    #3

    Exited with G3 determinism FAILED (model_identical=False json_identical=True) after all five arms were built and passed G1, G4 and G5.

    Exit code 1. Error: G3 determinism FAILED

    [G3] retraining e9-pl for determinism check
      G3: model_identical=False json_identical=True
    G3 determinism FAILED
    • arm_inputs/e9-sci.txt · Not redistributed: ask the Room owner, or rebuild with scripts/02_train_tokenizers.py from the training text · restricted · sha256 849d247145f5… · 1080627195 bytes · M420
    • arm_inputs/e9-sci-ws.txt · Not redistributed: ask the Room owner, or rebuild with scripts/02_train_tokenizers.py from the training text · restricted · sha256 849d247145f5… · 1080627195 bytes · M420
    • arm_inputs/e9-bal.txt · Not redistributed: ask the Room owner, or rebuild with scripts/02_train_tokenizers.py from the training text · restricted · sha256 9cff8d2fe22b… · 1081098259 bytes · M421
    • e9-pl/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/tokenizer.json · public · sha256 f22aca2a7ccc… · 3715831 bytes · M422
    • e9-pl/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
    • e9-pl/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
    • e9-pl/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/spm.vocab · public · sha256 25b8e27e518e… · 517037 bytes · M425
    • e9-pl/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/training_meta.json · public · sha256 a731cf8f0c35… · 1241 bytes · M426
    • e9-pl/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-pl/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 6d9e76290c03… · 562280 bytes · M427
    • e9-pl-fix/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/tokenizer.json · public · sha256 23b1d93a66e5… · 3715477 bytes · M428
    • e9-pl-fix/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
    • e9-pl-fix/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
    • e9-pl-fix/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/spm.vocab · public · sha256 59d0e6687fae… · 516998 bytes · M429
    • e9-pl-fix/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/training_meta.json · public · sha256 526a8f728d36… · 1300 bytes · M430
    • e9-pl-fix/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-pl-fix/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 0b01d5386a6e… · 562287 bytes · M431
    • e9-sci/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/tokenizer.json · public · sha256 f93babbc4752… · 3863491 bytes · M432
    • e9-sci/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
    • e9-sci/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
    • e9-sci/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/spm.vocab · public · sha256 90e84b0ccb97… · 498457 bytes · M433
    • e9-sci/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/training_meta.json · public · sha256 22f441a88f06… · 1808 bytes · M434
    • e9-sci/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-sci/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 735165bf143e… · 543740 bytes · M435
    • e9-sci-ws/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/tokenizer.json · public · sha256 2865e8236f0d… · 3870118 bytes · M436
    • e9-sci-ws/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
    • e9-sci-ws/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
    • e9-sci-ws/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/spm.vocab · public · sha256 de8b524b3bc1… · 498631 bytes · M437
    • e9-sci-ws/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/training_meta.json · public · sha256 62fa7f4df3b5… · 1813 bytes · M438
    • e9-sci-ws/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-sci-ws/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 66247c0392de… · 543920 bytes · M439
    • e9-bal/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/tokenizer.json · public · sha256 dd76a944cc77… · 3813566 bytes · M440
    • e9-bal/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
    • e9-bal/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
    • e9-bal/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/spm.vocab · public · sha256 b59020b8ab7d… · 491598 bytes · M441
    • e9-bal/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/training_meta.json · public · sha256 2b33dbbcc5bc… · 1982 bytes · M442
    • e9-bal/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-bal/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 2648ebb94571… · 536881 bytes · M443
    • spm_work/e9-pl-g3.model · Not redistributed (trainer work file): ask the Room owner, or rebuild with scripts/02_train_tokenizers.py --g3 · restricted · sha256 1ee3284a351d… · 562283 bytes · M444

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.