Finding

Sign in with GitHub
← Publications

Finding · P67 · Author-curated

Bielik-11B-v3.0-Instruct's arithmetic accuracy with space-grouped numbers minus its accuracy with no-break-space-grouped numbers is -0.0882 (95% interval -0.1035 to -0.0737; unadjusted bootstrap p-value 0 with 1,000 resamples).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Setting contrast

metric
Arithmetic probe accuracyQuantity compared.
subject
Bielik-11B-v3.0-InstructModel measured.
setting a
Space-grouped number formatFirst number format.
setting b
No-break-space-grouped number formatSecond number format.
value
-0.0882 difference in share of itemsContrast estimate.
interval low
-0.1035 difference in share of itemsLower bound of the 95% interval.
interval high
-0.0737 difference in share of itemsUpper bound of the 95% interval.
p bootstrap
0 p-valueUnadjusted bootstrap p-value.
scope
Items with at least four digits in addition, subtraction, sorting and unit conversion, with English and Polish instructions; comparison excluded as uninformative.Items the contrast is computed on.

Experimental provenance

Method and evaluation protocol
Paired item-level bootstrap over the probe items within each task, 1,000 resamples, seed 20260703; Holm adjustment over C1, C4', C2 for the PL model (space to no-break space) and C5; strata where both models' pooled accuracy lies between 5% and 95%.
Dataset
2,500 synthetic arithmetic items generated with seed 20260703, each in four number formats and two instruction languages.Version: unspecified · Access: public
Reported results
-0.0882 [-0.1035, -0.0737].
Uncertainty and replication
95% paired bootstrap interval over items, 1,000 resamples. A bootstrap p-value of 0 means no resample of 1,000 fell on the other side of zero.
Limitations
Outside the Holm family; the original report listed its p-value as Holm-adjusted. The two models differ by 20B tokens of continued pretraining and by post-training as well as by the tokenizer; answer-only probes do not measure chain-of-thought arithmetic; MLX bf16 numerics differ from the paper's evaluation stack.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Setting contrast

value is the subject's metric in setting_a minus its metric in setting_b over the same items in scope; interval_low and interval_high bound its 95% paired bootstrap interval; p_holm is its Holm-adjusted bootstrap p-value and p_bootstrap its unadjusted bootstrap p-value, where given.

Key setting_contrast · version 81582930-04ae-4afe-809e-f2d13cec4c18

Arithmetic probe accuracy

Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.

Key probe_accuracy · version 5b6b6982-53e3-4092-93d1-b3986055daf1

Space-grouped number format

Numbers with thousands grouped by spaces and a decimal comma, as in 1 234,56.

Key space_number_format · version 5b6b6982-53e3-4092-93d1-b3986055daf1

No-break-space-grouped number format

Numbers with thousands grouped by no-break spaces (U+00A0) and a decimal comma.

Key nbsp_number_format · version 1235918d-539a-46a7-85fe-2176fb7bcfee

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c