Finding

Sign in with GitHub
← Publications

Finding · P69 · Author-curated

Bielik-PL-11B-v3.0-Instruct's arithmetic accuracy minus Bielik-11B-v3.0-Instruct's, averaged over the four number formats, is -0.0299 (95% interval -0.0387 to -0.0215; Holm-adjusted bootstrap p-value 0 with 1,000 resamples).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Arithmetic probe accuracyQuantity compared.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
settings
bare, comma-grouped, space-grouped and no-break-space-grouped number formatsNumber formats averaged over.
value
-0.0299 difference in share of itemsDifference estimate.
interval low
-0.0387 difference in share of itemsLower bound of the 95% interval.
interval high
-0.0215 difference in share of itemsUpper bound of the 95% interval.
p holm
0 p-valueHolm-adjusted bootstrap p-value.
scope
Items with at least four digits in addition, subtraction, sorting and unit conversion, with English and Polish instructions; comparison excluded as uninformative.Items the difference is computed on.

Experimental provenance

Method and evaluation protocol
Paired item-level bootstrap over the probe items within each task, 1,000 resamples, seed 20260703; Holm adjustment over C1, C4', C2 for the PL model (space to no-break space) and C5; strata where both models' pooled accuracy lies between 5% and 95%.
Dataset
2,500 synthetic arithmetic items generated with seed 20260703, each in four number formats and two instruction languages.Version: unspecified · Access: public
Reported results
-0.0299 [-0.0387, -0.0215].
Uncertainty and replication
95% paired bootstrap interval over items, 1,000 resamples. A bootstrap p-value of 0 means no resample of 1,000 fell on the other side of zero.
Limitations
By the plan's attribution rule this is a transplant-pipeline effect, not a tokenizer effect. The two models differ by 20B tokens of continued pretraining and by post-training as well as by the tokenizer; answer-only probes do not measure chain-of-thought arithmetic; MLX bf16 numerics differ from the paper's evaluation stack. No difference excluding the no-break-space format was computed.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Arithmetic probe accuracy

Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.

Key probe_accuracy · version 5b6b6982-53e3-4092-93d1-b3986055daf1

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04