Finding

Sign in with GitHub
← Publications

Finding · P199 · Author-curated

On 295 of 300 LLMzSzŁ STEM questions with Polish chain-of-thought, the mean normalised pivot of Bielik-PL-11B-v3.0-Instruct minus that of Bielik-11B-v3.0-Instruct is -0.6854 (95% interval -0.6928 to -0.678; bootstrap p-value with 2,000 resamples, Holm-adjusted over the three benchmarks, 0).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Normalised pivotQuantity compared.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
benchmark
LLMzSzŁ STEM questionsBenchmark the questions come from.
value
-0.6854 normalised shareMean over questions of subject minus comparator.
interval low
-0.6928 normalised shareLower bound of the 95% paired bootstrap interval.
interval high
-0.678 normalised shareUpper bound of the 95% paired bootstrap interval.
p holm
0 p-valueTwo-sided bootstrap p-value, Holm-adjusted over the three benchmarks.
scope
295 of 300 LLMzSzŁ STEM questions usable in both models, answers with Polish chain-of-thoughtQuestions and answers compared.

Experimental provenance

Method and evaluation protocol
Teacher-forced forward pass of each answer through the MLX bf16 model with a custom loop over the 50 blocks; after each block the residual stream at generated positions is decoded through the final RMSNorm and output head with an fp32 softmax; masses on English- and Polish-labelled pieces are pooled over reasoning positions whose next token is language-labelled. Answers below 0.95 teacher-forcing agreement, with a failed round-trip or without language positions are excluded. Paired difference of normalised pivot restricted to one benchmark.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
-0.6854 [-0.6928, -0.678], Holm p 0 over three benchmarks.
Uncertainty and replication
95% paired item bootstrap interval within the benchmark (2000 resamples, seed 20260704); Holm adjustment over the three benchmarks. A p-value of 0 means that no bootstrap resample reached zero; with 2000 resamples the smallest non-zero value is 0.001.
Limitations
Logit lens, not a tuned lens: early-layer decodings are noisy. The vocabulary partitions label 16163 Polish and 1289 English pieces for APT4 and 1665 and 6116 for the Mistral-derived tokenizer, and were built from re-serialised MLX tokenizer files without an equivalence check. The design was written after the lens run started. The Polish-model lens file combines records from two invocations separated by an unrecorded rewrite of the lens script.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

LLMzSzŁ STEM questions

300 multiple-choice questions in mathematics, physics, science and biology from the LLMzSzŁ test split, excluding vocational examinations and questions that refer to figures or tables.

Key llmzszl_stem_questions · version 4c39157e-9b47-4096-b640-a79bf90a103d

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Normalised pivot

For one answer, the band English share minus the Polish anchor, divided by the English anchor minus the Polish anchor, where the anchors are the model's pooled layer-50 English shares on its answers with Polish and with English chain-of-thought to the questions of the same benchmark; 0 is the model's overt-Polish level and 1 its overt-English level.

Key normalised_pivot · version e142a14d-192a-4ae7-ad70-c2300ef06faf

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

refines

e142a14d-192a-4ae7-ad70-c2300ef06faf

The pooled difference, split by benchmark.