Arithmetic probe accuracy
Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.
Hypothesis
Sign in with GitHubHypothesis · H5 · Author-curated prediction
Answer-only arithmetic probes in English and Polish instructions under greedy decoding, pooled across tasks whose accuracy is between 5% and 95% for both models, on items with at least four digits.
cited claim · premise
Beyond vocabulary size, the handling of digits, punctuation and special characters can influence both token efficiency and downstream generation quality.Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.
Loading research…
Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.
The subject's metric minus the comparator's metric in setting_a, minus the same difference in setting_b, has the sign given in sign.
Numbers with thousands grouped by commas and a decimal point, as in 1,234.56.
Numbers with thousands grouped by spaces and a decimal comma, as in 1 234,56.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
{
"wording": "Bielik-PL-11B-v3.0-Instruct's arithmetic accuracy deficit against Bielik-11B-v3.0-Instruct is larger with comma-grouped numbers than with space-grouped numbers.",
"predicate": {
"type": "concept",
"key": "interaction_sign"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "probe_accuracy"
}
},
{
"role": "subject",
"definition": "Model whose deficit is measured.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "comparator",
"definition": "Model the deficit is measured against.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "setting_a",
"definition": "Number format with the predicted larger deficit.",
"value": {
"type": "concept",
"key": "comma_number_format"
}
},
{
"role": "setting_b",
"definition": "Number format compared with.",
"value": {
"type": "concept",
"key": "space_number_format"
}
},
{
"role": "sign",
"definition": "Predicted sign of the interaction.",
"value": {
"type": "text",
"value": "negative"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →