Finding

Sign in with GitHub
← Publications

Finding · P232 · Author-curated · Superseded

Training the Polish-only 32k tokenizer a second time from the same input gives identical model content once the stored output path is cleared and a byte-identical tokenizer.json; the raw model files differ.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Notices · Superseded · this exact version stays citable

Structured assertion

Relation: Retrain identity

subject
Polish-only 32k tokenizerTokenizer trained twice.
content identical
trueModel content identical with the output path cleared.
raw identical
falseRaw model files byte-identical.
converted identical
trueConverted tokenizer.json byte-identical.

Experimental provenance

Method and evaluation protocol
Retrain the arm from its frozen input into a second output path, compare the serialised models with trainer_spec.model_prefix cleared, the raw files and the converted tokenizer.json.
Dataset
The Polish-only arm's 1 GiB training input.Version: unspecified · Access: restricted
Reported results
Model content identical true; raw bytes identical false; tokenizer.json identical true. With trainer_spec.input also cleared, the retrained model equals the shipped Polish-only spm.model.
Uncertainty and replication
One retrain of one arm.
Evidence references
tokenizers_e9_manifest.json · unavailableArchived private repository, experiments/E9/results/tokenizers_e9_manifest.json at commit e4f0cad; the public copy lists the five scrubbed spm.model digestsmetrics.json · publichttps://github.com/stw2/tokenizer-science-tax/blob/c5299916db7f792a209ff2944e66c673c3ebb817/experiments/E09-vocabulary-allocation/results/metrics.json
Limitations
The first comparison, on raw bytes, failed; the rule was amended afterwards and the passing comparison comes from an unlogged re-run. The byte-identical atlas re-run the gate also requires has no record.

Author’s note

Values from G3_determinism in the tokenizers manifest; the last sentence of the results was checked during the port.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Retrain identity

Training the subject twice from the same input gives content_identical for the model content with the stored output path cleared, raw_identical for the raw model files and converted_identical for the converted tokenizer.json.

Key retrain_identity · version 2a78e81e-96a3-4141-aeb3-c5453c325fb0

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd