Finding

Sign in with GitHub
← Publications

Finding · P60 · Author-curated

On arithmetic probes with bare numbers, pooled over five tasks and two instruction languages, Bielik-PL-11B-v3.0-Instruct answers 0.8582 of 5,000 items correctly and Bielik-11B-v3.0-Instruct 0.8428.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Arithmetic probe accuracyQuantity compared.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
subject value
0.8582 share of itemsValue of the metric for the subject.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
comparator value
0.8428 share of itemsValue of the metric for the comparator.
setting
Bare number formatNumber format of the items.
scope
5,000 items per model: 500 per task for addition, subtraction, comparison, sorting and unit conversion, each with English and Polish instructions; all digit lengthsItems pooled.

Experimental provenance

Method and evaluation protocol
Greedy generation with at most 256 tokens; dual-locale grading of the last number in the answer; per-cell counts pooled over tasks and instruction languages.
Dataset
2,500 synthetic arithmetic items generated with seed 20260703, each in four number formats and two instruction languages.Version: unspecified · Access: public
Reported results
0.8582 against 0.8428; truncated answers 0 and 0; answers ambiguous between the English and Polish readings 4 and 3; answers without a number 0 and 0.
Uncertainty and replication
Point values over 5,000 items per model; per-cell Wilson intervals are in results/probe_results.json.
Limitations
The two models differ by 20B tokens of continued pretraining and by post-training as well as by the tokenizer; answer-only probes do not measure chain-of-thought arithmetic; MLX bf16 numerics differ from the paper's evaluation stack.

Author’s note

Pooled by scripts/09_metrics.py from the per-cell counts in results/probe_results.json.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Bare number format

Numbers without thousands grouping and with a decimal point, as in 1234.56.

Key bare_number_format · version a1731117-7a38-4a69-aed5-baaadefc0608

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Arithmetic probe accuracy

Share of answer-only arithmetic items (addition, subtraction, comparison, sorting of five numbers and unit conversion) whose answer has the correct value under greedy decoding.

Key probe_accuracy · version 5b6b6982-53e3-4092-93d1-b3986055daf1

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04