Finding

Sign in with GitHub
← Publications

Finding · P226 · Author-curated

The Polish-only 32k tokenizer's fertility tax against the Mistral-derived tokenizer is 0.9727, 0.9882, 0.9965 and 0.9713 times APT4's on English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, and 0.9586, 0.9577 and 1.0065 times APT4's on Polish PES examination questions, Polish Wikipedia science articles and Polish reviews.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Ratio per corpus

metric
Fertility taxQuantity compared.
subject
Polish-only 32k tokenizerTokenizer measured.
comparator
APT4 tokenizerTokenizer compared with.
baseline
Mistral-derived tokenizerTokenizer both taxes are measured against.
ratio english arxiv abstracts
0.9727 ratioRatio on English arXiv abstracts.
ratio english latex methods
0.9882 ratioRatio on English LaTeX method sections.
ratio english python code
0.9965 ratioRatio on English Python code.
ratio english gsm8k problems
0.9713 ratioRatio on GSM8K problems.
ratio polish pes questions
0.9586 ratioRatio on Polish PES examination questions.
ratio polish wikipedia science
0.9577 ratioRatio on Polish Wikipedia science articles.
ratio polish reviews
1.0065 ratioRatio on Polish reviews.

Experimental provenance

Method and evaluation protocol
Divide the Polish-only arm's token ratio to the Mistral-derived tokenizer by APT4's on each corpus.
Dataset
Seven English and Polish corpora of about 120k words each; four are restricted.Version: unspecified · Access: restricted
Reported results
en-arxiv-abstracts 0.9727; en-latex-methods 0.9882; en-python-code 0.9965; en-gsm8k 0.9713; pl-science-pes 0.9586; pl-wiki-science 0.9577; pl-informal 1.0065
Uncertainty and replication
Point ratios; intervals of each tax are in results/atlas.json.
Limitations
APT4's training corpus is unknown, so differences mix recipe and data. The original report's prose said within 1 to 3% on every corpus; the Polish PES and Wikipedia ratios differ by more than 4% (report.md, Corrections, 14 September 2026).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Ratio per corpus

Each role ratio_<corpus> is the subject's metric divided by the comparator's metric on that corpus, both against the baseline tokenizer.

Key ratio_per_corpus · version 23b1d16c-84ca-464f-a2d6-421d0f9e7f26

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Fertility tax

Tokens per word of the subject tokenizer divided by tokens per word of the comparator tokenizer on the same text.

Key fertility_tax · version 1de64025-ce26-4aa5-a17b-0ebb58006700

Polish-only 32k tokenizer

SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.

Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd

Exact references

related

3d1e1910-f471-4b1a-be65-5e2db901f9a0

APT4's taxes on the same English corpora.