Finding

Sign in with GitHub
← Publications

Finding · P198 · Author-curated

On 248 of 250 GSM8K-PL problems with Polish chain-of-thought, the mean normalised pivot of Bielik-PL-11B-v3.0-Instruct minus that of Bielik-11B-v3.0-Instruct is -0.5882 (95% interval -0.5961 to -0.5806; bootstrap p-value with 2,000 resamples, Holm-adjusted over the three benchmarks, 0).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Normalised pivotQuantity compared.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
benchmark
GSM8K-PL problemsBenchmark the questions come from.
value
-0.5882 normalised shareMean over questions of subject minus comparator.
interval low
-0.5961 normalised shareLower bound of the 95% paired bootstrap interval.
interval high
-0.5806 normalised shareUpper bound of the 95% paired bootstrap interval.
p holm
0 p-valueTwo-sided bootstrap p-value, Holm-adjusted over the three benchmarks.
scope
248 of 250 GSM8K-PL problems usable in both models, answers with Polish chain-of-thoughtQuestions and answers compared.

Experimental provenance

Method and evaluation protocol
Teacher-forced forward pass of each answer through the MLX bf16 model with a custom loop over the 50 blocks; after each block the residual stream at generated positions is decoded through the final RMSNorm and output head with an fp32 softmax; masses on English- and Polish-labelled pieces are pooled over reasoning positions whose next token is language-labelled. Answers below 0.95 teacher-forcing agreement, with a failed round-trip or without language positions are excluded. Paired difference of normalised pivot restricted to one benchmark.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
-0.5882 [-0.5961, -0.5806], Holm p 0 over three benchmarks.
Uncertainty and replication
95% paired item bootstrap interval within the benchmark (2000 resamples, seed 20260704); Holm adjustment over the three benchmarks. A p-value of 0 means that no bootstrap resample reached zero; with 2000 resamples the smallest non-zero value is 0.001.
Limitations
Logit lens, not a tuned lens: early-layer decodings are noisy. The vocabulary partitions label 16163 Polish and 1289 English pieces for APT4 and 1665 and 6116 for the Mistral-derived tokenizer, and were built from re-serialised MLX tokenizer files without an equivalence check. The design was written after the lens run started. The Polish-model lens file combines records from two invocations separated by an unrecorded rewrite of the lens script.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

GSM8K-PL problems

250 GSM8K test problems in the Polish machine translation of gsm8kx, stratified by the digit length of the gold answer, with openai/gsm8k gold answers.

Key gsm8k_pl_problems · version 080dd7b3-7a21-419e-bf4e-c3770f5cdde6

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Normalised pivot

For one answer, the band English share minus the Polish anchor, divided by the English anchor minus the Polish anchor, where the anchors are the model's pooled layer-50 English shares on its answers with Polish and with English chain-of-thought to the questions of the same benchmark; 0 is the model's overt-Polish level and 1 its overt-English level.

Key normalised_pivot · version e142a14d-192a-4ae7-ad70-c2300ef06faf

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

refines

e142a14d-192a-4ae7-ad70-c2300ef06faf

The pooled difference, split by benchmark.