Sign of an interaction
The subject's metric minus the comparator's metric in setting_a, minus the same difference in setting_b, has the sign given in sign.
Hypothesis
Sign in with GitHubHypothesis · H6 · Author-curated prediction
The same probe items with the same visible numbers, differing only in the grouping character; items with at least four digits.
hypothesis · premise
APT4 uses more tokens per number than the Mistral-derived tokenizer on numbers grouped with no-break spaces.APT4 splits the no-break space into more tokens than the Mistral-derived tokenizer.
cited claim · premise
Beyond vocabulary size, the handling of digits, punctuation and special characters can influence both token efficiency and downstream generation quality.Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.
Loading research…
The subject's metric minus the comparator's metric in setting_a, minus the same difference in setting_b, has the sign given in sign.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.
Numbers with thousands grouped by spaces and a decimal comma, as in 1 234,56.
Numbers with thousands grouped by no-break spaces (U+00A0) and a decimal comma.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
{
"wording": "Replacing space grouping with no-break-space grouping lowers the arithmetic accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct.",
"predicate": {
"type": "concept_ref",
"versionId": "5b6b6982-53e3-4092-93d1-b3986055daf1",
"key": "interaction_sign"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept_ref",
"versionId": "5b6b6982-53e3-4092-93d1-b3986055daf1",
"key": "probe_accuracy"
}
},
{
"role": "subject",
"definition": "Model whose tokenizer splits the no-break space into bytes.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "comparator",
"definition": "Model whose tokenizer has a no-break-space token.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "setting_a",
"definition": "Number format before the replacement.",
"value": {
"type": "concept_ref",
"versionId": "5b6b6982-53e3-4092-93d1-b3986055daf1",
"key": "space_number_format"
}
},
{
"role": "setting_b",
"definition": "Number format after the replacement.",
"value": {
"type": "concept_ref",
"versionId": "1235918d-539a-46a7-85fe-2176fb7bcfee",
"key": "nbsp_number_format"
}
},
{
"role": "sign",
"definition": "Predicted sign of the interaction.",
"value": {
"type": "text",
"value": "positive"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →