Finding

Sign in with GitHub
← Publications

Finding · P228 · Author-curated

In a 32,768-token window, the science-slice 32k tokenizer fits 0.8258, the science-slice 32k tokenizer with whitespace pieces 0.9188, the balanced 32k tokenizer 0.8635, the Polish-only 32k tokenizer 0.6546 and APT4 0.6442 of the English-science characters the Mistral-derived tokenizer fits.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Values per tokenizer

metric
Effective context capacityQuantity measured.
reference
Mistral-derived tokenizerTokenizer each value is relative to.
value science slice
0.8258 ratio to referenceValue for the science-slice 32k tokenizer.
value science slice whitespace
0.9188 ratio to referenceValue for the science-slice 32k tokenizer with whitespace pieces.
value balanced
0.8635 ratio to referenceValue for the balanced 32k tokenizer.
value polish only
0.6546 ratio to referenceValue for the Polish-only 32k tokenizer.
value apt4
0.6442 ratio to referenceValue for APT4.
evaluation text
English arXiv abstracts, LaTeX method sections, Python code and GSM8K problems, pooledCorpora measured.
language
EnglishLanguage of the corpora.

Experimental provenance

Method and evaluation protocol
Characters of the four pooled English corpora per token times 32,768, divided by the same quantity for the Mistral-derived tokenizer.
Dataset
Four English corpora of about 120k words each: arXiv abstracts, LaTeX method sections, Python code, GSM8K problems; two are restricted.Version: unspecified · Access: restricted
Reported results
e9-sci 0.8258 (93,867 characters); e9-sci-ws 0.9188 (104,432 characters); e9-bal 0.8635 (98,148 characters); e9-pl 0.6546 (74,401 characters); apt4 0.6442 (73,225 characters); Mistral-derived 113,664 characters.
Uncertainty and replication
Point values from total characters and tokens; no interval.
Limitations
Window arithmetic only; whether a model uses the window better is not measured.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Values per tokenizer

The metric, relative to the reference tokenizer on the evaluation text, takes value_science_slice for the science-slice 32k tokenizer, value_science_slice_whitespace for the science-slice 32k tokenizer with whitespace pieces, value_balanced for the balanced 32k tokenizer, value_polish_only for the Polish-only 32k tokenizer and value_apt4 for APT4.

Key values_per_tokenizer · version c212f70e-53ad-49af-8989-afba67519c2f

Effective context capacity

Amount of text that fits in a model's context window.

Key effective_context_capacity · version 9ac91b29-87d3-48c1-a32f-95da879874b3

Mistral-derived tokenizer

Tokenizer of the original Bielik v3 models, derived from Mistral's.

Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7

Exact references

related

9ac91b29-87d3-48c1-a32f-95da879874b3

The same arithmetic on Polish text, where the Polish-optimised tokenizer gains.

related

538541db-aa28-4110-9de1-f59a7a701952

Measured character ceilings of the models with APT4 and with the Mistral-derived tokenizer at a 32,704-token prompt limit on English scientific documents.

related

9d6d1a8a-1772-421f-962b-67f9475cd006

Measured effective context in characters of the same two models on English scientific documents.