Accounts for
The difference in factor between subject_tokenizer and comparator_tokenizer causes the subject model's lower score than the comparator model's on the benchmark.
Hypothesis
Sign in with GitHubHypothesis · H2 · Author-curated prediction
The 11B instruction-tuned Bielik v3 pair; tokenization of digits and number separators; GSM8K as reported in arXiv:2604.10799v1.
cited claim · premise
On GSM8K, Bielik-PL-11B-v3.0-Instruct scores 80.97 and Bielik-11B-v3.0-Instruct 85.60.The score difference the mechanism is proposed for.
cited claim · premise
Beyond vocabulary size, the handling of digits, punctuation and special characters can influence both token efficiency and downstream generation quality.Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 1.3640 on GSM8K problems (95% interval 1.3616 to 1.3665), below its 1.5466 on the English preamble.APT4's fertility tax on GSM8K questions is below its tax on the English preamble, so question length is not the proposed cause.
cited claim · related
English-language capabilities of the Bielik v3 PL models remain largely intact.The GSM8K difference is the one disputing this claim.
Loading research…
The difference in factor between subject_tokenizer and comparator_tokenizer causes the subject model's lower score than the comparator model's on the benchmark.
Mathematical reasoning task of the English Open LLM Leaderboard.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
Tokenizer of the original Bielik v3 models, derived from Mistral's.
How a tokenizer segments digits, punctuation and special characters.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Polish-optimised tokenizer of the Bielik v3 PL models.
{
"wording": "The difference between APT4's and the Mistral-derived tokenizer's handling of digits accounts for the lower GSM8K score of Bielik-PL-11B-v3.0-Instruct than of Bielik-11B-v3.0-Instruct.",
"predicate": {
"type": "concept",
"key": "accounts_for"
},
"roles": [
{
"role": "factor",
"definition": "Tokenizer property whose difference is the proposed cause.",
"value": {
"type": "concept_ref",
"versionId": "85dfefb0-eb84-4deb-b93c-7b494a10b283",
"key": "digit_tokenization_policy"
}
},
{
"role": "subject_tokenizer",
"definition": "Tokenizer of the lower-scoring model.",
"value": {
"type": "concept_ref",
"versionId": "d497f94d-5373-4652-887c-55c001b6472c",
"key": "apt4"
}
},
{
"role": "comparator_tokenizer",
"definition": "Tokenizer of the higher-scoring model.",
"value": {
"type": "concept_ref",
"versionId": "bd3d8329-e626-4805-9fff-f83f146769a7",
"key": "mistral_tokenizer"
}
},
{
"role": "subject",
"definition": "Model with the lower score.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "comparator",
"definition": "Model with the higher score.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "benchmark",
"definition": "Benchmark of the score difference.",
"value": {
"type": "concept_ref",
"versionId": "36da68d9-97cf-4575-ad42-ed56e0e895e9",
"key": "gsm8k"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →