Execution attempt · A85 · planned
Trains the five 32k arms and converts each to a fast tokenizer; gates G1, G4 and G5 pass for every arm. Then retrains the Polish-only arm for the determinism gate G3, which fails: the retrained model file is not byte-identical, tokenizer.json is. The arm inputs of the two Polish-only arms are byte-identical to pl.txt and are not listed. The G3 comparison code of this invocation predates Amendment 2 and was not kept; the pinned commit holds the amended code.
Pinned source and configuration
https://github.com/stw2/tokenizer-science-tax @ f52d5afd8e13cbace7a438d4fd653deeb64de610
Reference checked 2026-09-14 13:15 UTC. Later commits, branches or plan changes do not retarget this attempt.
- Command
- python scripts/02_train_tokenizers.py --g3
- Working directory
- experiments/E09-vocabulary-allocation
- Configuration paths
- None
- Parameters
- vocabulary 32,000; BPE; character coverage 0.9995; identity normalisation; byte fallback; split digits; maximum piece length 16; no input shuffling; separator pieces as user-defined symbols in four arms; whitespace-only pieces in one arm; num_threads 8
- Environment
- 3.12.13; SentencePiece BPE trainer and Hugging Face tokenizers on CPU; no model weights; Apple silicon, 128 GB unified memory
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- pl.txt · training data · M351 Polish FineWeb2-HQ training text (pl.txt) · dataset · Ask the reporter
- en_general.txt · training data · M418 en_general.txt · raw output · Ask the reporter
- sci_latex.txt · training data · M415 sci_latex.txt · raw output · Ask the reporter
- sci_python.txt · training data · M416 sci_python.txt · raw output · Ask the reporter
- sci_mathweb.txt · training data · M417 sci_mathweb.txt · raw output · Ask the reporter
- tokenizer apt4 · reference tokenizer · M3 APT4 tokenizer (Bielik-PL-11B-v3.0-Instruct) · tokenizer · Ask the reporter
Delivered events
Registered
#1Trains the five 32k arms and converts each to a fast tokenizer; gates G1, G4 and G5 pass for every arm. Then retrains the Polish-only arm for the determinism gate G3, which fails: the retrained model file is not byte-identical, tokenizer.json is. The arm inputs of the two Polish-only arms are byte-identical to pl.txt and are not listed. The G3 comparison code of this invocation predates Amendment 2 and was not kept; the pinned commit holds the amended code.
Started
#2Started.
Failed
#3Exited with G3 determinism FAILED (model_identical=False json_identical=True) after all five arms were built and passed G1, G4 and G5.
Exit code 1. Error: G3 determinism FAILED
[G3] retraining e9-pl for determinism check G3: model_identical=False json_identical=True G3 determinism FAILED
- arm_inputs/e9-sci.txt · Not redistributed: ask the Room owner, or rebuild with scripts/02_train_tokenizers.py from the training text · restricted · sha256 849d247145f5… · 1080627195 bytes · M420
- arm_inputs/e9-sci-ws.txt · Not redistributed: ask the Room owner, or rebuild with scripts/02_train_tokenizers.py from the training text · restricted · sha256 849d247145f5… · 1080627195 bytes · M420
- arm_inputs/e9-bal.txt · Not redistributed: ask the Room owner, or rebuild with scripts/02_train_tokenizers.py from the training text · restricted · sha256 9cff8d2fe22b… · 1081098259 bytes · M421
- e9-pl/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/tokenizer.json · public · sha256 f22aca2a7ccc… · 3715831 bytes · M422
- e9-pl/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
- e9-pl/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
- e9-pl/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/spm.vocab · public · sha256 25b8e27e518e… · 517037 bytes · M425
- e9-pl/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl/training_meta.json · public · sha256 a731cf8f0c35… · 1241 bytes · M426
- e9-pl/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-pl/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 6d9e76290c03… · 562280 bytes · M427
- e9-pl-fix/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/tokenizer.json · public · sha256 23b1d93a66e5… · 3715477 bytes · M428
- e9-pl-fix/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
- e9-pl-fix/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
- e9-pl-fix/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/spm.vocab · public · sha256 59d0e6687fae… · 516998 bytes · M429
- e9-pl-fix/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-pl-fix/training_meta.json · public · sha256 526a8f728d36… · 1300 bytes · M430
- e9-pl-fix/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-pl-fix/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 0b01d5386a6e… · 562287 bytes · M431
- e9-sci/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/tokenizer.json · public · sha256 f93babbc4752… · 3863491 bytes · M432
- e9-sci/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
- e9-sci/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
- e9-sci/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/spm.vocab · public · sha256 90e84b0ccb97… · 498457 bytes · M433
- e9-sci/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci/training_meta.json · public · sha256 22f441a88f06… · 1808 bytes · M434
- e9-sci/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-sci/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 735165bf143e… · 543740 bytes · M435
- e9-sci-ws/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/tokenizer.json · public · sha256 2865e8236f0d… · 3870118 bytes · M436
- e9-sci-ws/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
- e9-sci-ws/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
- e9-sci-ws/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/spm.vocab · public · sha256 de8b524b3bc1… · 498631 bytes · M437
- e9-sci-ws/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-sci-ws/training_meta.json · public · sha256 62fa7f4df3b5… · 1813 bytes · M438
- e9-sci-ws/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-sci-ws/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 66247c0392de… · 543920 bytes · M439
- e9-bal/tokenizer.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/tokenizer.json · public · sha256 dd76a944cc77… · 3813566 bytes · M440
- e9-bal/tokenizer_config.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/tokenizer_config.json · public · sha256 fa53599038cf… · 216 bytes · M423
- e9-bal/special_tokens_map.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/special_tokens_map.json · public · sha256 96bdbb8504d9… · 72 bytes · M424
- e9-bal/spm.vocab · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/spm.vocab · public · sha256 b59020b8ab7d… · 491598 bytes · M441
- e9-bal/training_meta.json · https://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/data/tokenizers/e9-bal/training_meta.json · public · sha256 2b33dbbcc5bc… · 1982 bytes · M442
- e9-bal/spm.model · Archived private repository, experiments/E9/data/tokenizers/e9-bal/spm.model at commit e4f0cad; the public copy clears trainer_spec.input and trainer_spec.model_prefix, which held absolute paths · unavailable · sha256 2648ebb94571… · 536881 bytes · M443
- spm_work/e9-pl-g3.model · Not redistributed (trainer work file): ask the Room owner, or rebuild with scripts/02_train_tokenizers.py --g3 · restricted · sha256 1ee3284a351d… · 562283 bytes · M444
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.