Finding

Sign in with GitHub
← Publications

Finding · P68 · Author-curated

For both models, arithmetic accuracy with bare numbers minus accuracy with comma-grouped or space-grouped numbers ranges from 0.1651 to 0.1914 over 4 contrasts, with no 95% interval bound below 0.145.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Range of setting contrasts

metric
Arithmetic probe accuracyQuantity compared.
subjects
Bielik-PL-11B-v3.0-Instruct; Bielik-11B-v3.0-InstructModels measured.
setting a
Bare number formatNumber format subtracted from.
settings b
comma-grouped number format; space-grouped number formatNumber formats subtracted.
value min
0.1651 difference in share of itemsSmallest contrast.
value max
0.1914 difference in share of itemsLargest contrast.
interval low min
0.145 difference in share of itemsLowest lower bound of the 95% intervals.
count
4 contrastsContrasts.
scope
Items with at least four digits in addition, subtraction, sorting and unit conversion, with English and Polish instructions; comparison excluded as uninformative.Items the contrasts are computed on.

Experimental provenance

Method and evaluation protocol
Paired item-level bootstrap over the probe items within each task, 1,000 resamples, seed 20260703; Holm adjustment over C1, C4', C2 for the PL model (space to no-break space) and C5; strata where both models' pooled accuracy lies between 5% and 95%.
Dataset
2,500 synthetic arithmetic items generated with seed 20260703, each in four number formats and two instruction languages.Version: unspecified · Access: public
Reported results
Bielik-PL-11B-v3.0-Instruct bare minus comma-grouped: 0.1821 [0.1642, 0.1999]; Bielik-PL-11B-v3.0-Instruct bare minus space-grouped: 0.1914 [0.1701, 0.2125]; Bielik-11B-v3.0-Instruct bare minus comma-grouped: 0.1651 [0.145, 0.1842]; Bielik-11B-v3.0-Instruct bare minus space-grouped: 0.1779 [0.1566, 0.1985].
Uncertainty and replication
95% paired bootstrap intervals over items, 1,000 resamples; outside the Holm family.
Limitations
Grouped renderings apply only to integer parts of at least four digits, so the contrasts compare the same items with and without separators. The two models differ by 20B tokens of continued pretraining and by post-training as well as by the tokenizer; answer-only probes do not measure chain-of-thought arithmetic; MLX bf16 numerics differ from the paper's evaluation stack.

Author’s note

Range computed by scripts/09_metrics.py from the four C2 contrasts in results/probe_results.json.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Range of setting contrasts

Across count contrasts of each subject's metric in setting_a minus its metric in each of settings_b, over the items in scope, the values range from value_min to value_max and the lowest lower bound of their 95% paired bootstrap intervals is interval_low_min.

Key setting_contrast_range · version 064365b7-15d9-4516-89b7-51680b013cf7

Arithmetic probe accuracy

Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.

Key probe_accuracy · version 5b6b6982-53e3-4092-93d1-b3986055daf1

Bare number format

Numbers without thousands grouping and with a decimal point, as in 1234.56.

Key bare_number_format · version a1731117-7a38-4a69-aed5-baaadefc0608