Finding

Sign in with GitHub
← Publications

Finding · P230 · Author-curated

APT4's pooled English-science excess fertility tax against the Mistral-derived tokenizer is 0.5523 (95% interval 0.5377 to 0.5656) and APT3's 0.5007 (0.4827 to 0.5174).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

subject
APT4 tokenizerTokenizer measured.
subject value
0.5523 ratioValue of the metric for the subject.
comparator
APT3 tokenizerTokenizer compared with.
comparator value
0.5007 ratioValue of the metric for the comparator.
evaluation text
English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, pooledCorpora measured.
language
EnglishLanguage of the corpora.
setting
excess tax against the Mistral-derived tokenizerTokenizer both values are measured against.

Experimental provenance

Method and evaluation protocol
Sum each tokenizer's tokens over the four English corpora, divide by the Mistral-derived tokenizer's sum and subtract 1.
Dataset
Four English corpora of about 120k words each: arXiv abstracts, LaTeX method sections, Python code, GSM8K problems; two are restricted.Version: unspecified · Access: restricted
Reported results
APT4 0.5523 [0.5377, 0.5656]; APT3 0.5007 [0.4827, 0.5174]; Polish-only arm 0.5277 [0.5131, 0.5410].
Uncertainty and replication
95% percentile interval over 1000 document resamples within each corpus, paired across tokenizers (seed 20260703); tokenizer training is deterministic and was run once per arm.
Limitations
Each corpus is one sample of its domain.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

APT3 tokenizer

Polish tokenizer introduced with the Polish APT3 model.

Key apt3 · version a59f02ac-4303-44b5-a5b0-5decab214308

Pooled English-science excess fertility tax

Total tokens of a tokenizer over four English corpora (arXiv abstracts, LaTeX method sections, Python code, GSM8K problems) divided by the total tokens of the baseline tokenizer over the same corpora, minus 1.

Key pooled_excess_science_tax · version d1a46254-542b-43a9-b64b-3536a37e2efd

APT4 tokenizer

Polish-optimised tokenizer of the Bielik v3 PL models.

Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c

Exact references

extends

d12932ff-15a4-4030-ba68-04695fc77ae9

The same ordering on the four English corpora pooled.

related

e150c2d5-b823-43d9-9440-5f6428448be1

APT4 extends APT3's design; its pooled English-science excess tax is the higher of the two.