Ratio per corpus
Each role ratio_<corpus> is the subject's metric divided by the comparator's metric on that corpus, both against the baseline tokenizer.
Key ratio_per_corpus · version 23b1d16c-84ca-464f-a2d6-421d0f9e7f26
Finding
Sign in with GitHubFinding · P226 · Author-curated
Relation: Ratio per corpus
Reuse the defining version and key when the meaning fits your assertion.
Each role ratio_<corpus> is the subject's metric divided by the comparator's metric on that corpus, both against the baseline tokenizer.
Key ratio_per_corpus · version 23b1d16c-84ca-464f-a2d6-421d0f9e7f26
Tokenizer of the original Bielik v3 models, derived from Mistral's.
Key mistral_tokenizer · version bd3d8329-e626-4805-9fff-f83f146769a7
Polish-optimised tokenizer of the Bielik v3 PL models.
Key apt4 · version d497f94d-5373-4652-887c-55c001b6472c
Tokens per word of the subject tokenizer divided by tokens per word of the comparator tokenizer on the same text.
Key fertility_tax · version 1de64025-ce26-4aa5-a17b-0ebb58006700
SentencePiece BPE tokenizer with a 32,000-token vocabulary trained on 1 GiB of Polish web text from FineWeb2-HQ, with identity normalisation, a prepended word-boundary marker, byte fallback, digits split into single characters, pieces of at most 16 characters, no whitespace-only pieces and APT4's special-token and byte-piece id layout.
Key polish_only_32k_tokenizer · version d1a46254-542b-43a9-b64b-3536a37e2efd
Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.