Finding

Sign in with GitHub
← Publications

Finding · P195 · Author-curated

On 766 of 800 Polish STEM questions with Polish chain-of-thought, the mean pivot excess of Bielik-PL-11B-v3.0-Instruct is 0.1801, with a 95% bootstrap interval from 0.1731 to 0.187 (bootstrap p-value 0 with 2,000 resamples, Holm-adjusted 0).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Estimate with interval

metric
Pivot excessQuantity estimated.
subject
Bielik-PL-11B-v3.0-InstructModel measured.
evaluation items
Polish STEM question setQuestion set.
items
766 questionsQuestions used.
value
0.1801 shareMean over the questions used.
interval low
0.1731 shareLower bound of the 95% bootstrap interval.
interval high
0.187 shareUpper bound of the 95% bootstrap interval.
p value
0 p-valueTwo-sided bootstrap p-value.
setting
teacher-forced logit lens of answers with Polish chain-of-thought; two-sided stratified paired item bootstrap, 2000 resamplesMeasurement and test.
holm p
0 p-valuep_value after Holm correction over tests A1, A2 and B.

Experimental provenance

Method and evaluation protocol
Teacher-forced forward pass of each answer through the MLX bf16 model with a custom loop over the 50 blocks; after each block the residual stream at generated positions is decoded through the final RMSNorm and output head with an fp32 softmax; masses on English- and Polish-labelled pieces are pooled over reasoning positions whose next token is language-labelled. Answers below 0.95 teacher-forcing agreement, with a failed round-trip or without language positions are excluded. Per question, band English share minus the layer-50 share; mean over the questions usable in both models.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
Pivot excess 0.1801 [0.1731, 0.187], bootstrap p-value 0 with 2,000 resamples, Holm-adjusted 0; mean band English share 0.258, mean layer-50 English share 0.0779.
Uncertainty and replication
95% stratified paired item bootstrap interval (2000 resamples, strata benchmark, seed 20260704); Holm adjustment over A1, A2 and B. A p-value of 0 means that no bootstrap resample reached zero; with 2000 resamples the smallest non-zero value is 0.001. Sensitivity: truncation-excluded 0.1809 [0.174, 0.1878] (741 items); compliance-restricted 0.1818 [0.1746, 0.1886] (733 items); agreement at least 0.98 0.1663 [0.1568, 0.1751] (498 items); band 13 to 25 0.0442 [0.0375, 0.0505] (766 items); band 30 to 40 0.1588 [0.1519, 0.1655] (766 items); pooled ratio 0.182 [0.1749, 0.189] (766 items).
Limitations
Logit lens, not a tuned lens: early-layer decodings are noisy. The vocabulary partitions label 16163 Polish and 1289 English pieces for APT4 and 1665 and 6116 for the Mistral-derived tokenizer, and were built from re-serialised MLX tokenizer files without an equivalence check. The design was written after the lens run started. The Polish-model lens file combines records from two invocations separated by an unrecorded rewrite of the lens script.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Estimate with interval

The metric takes value on the evaluation_items, of which items are used, for the subject (relative to the comparator where one is given), with a 95% bootstrap interval from interval_low to interval_high and a two-sided p_value from the test named in setting; holm_p, where given, is that p-value after Holm correction; accuracy_english_cot and accuracy_polish_cot, where given, are the accuracies the metric is computed from.

Key estimate_with_interval · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Polish STEM question set

800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.

Key polish_stem_questions · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Pivot excess

The band English share of an answer minus its logit-lens English share at layer 50.

Key pivot_excess · version 25a042fa-58a1-4b0a-8984-73e43685a6e0

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04