Finding

Sign in with GitHub
← Publications

Finding · P197 · Author-curated

On 766 of 800 Polish STEM questions with Polish chain-of-thought, the mean normalised pivot is 0.2128 for Bielik-PL-11B-v3.0-Instruct and 0.9838 for Bielik-11B-v3.0-Instruct; the paired difference is -0.771 (95% interval -0.7764 to -0.7655; Holm-adjusted bootstrap p-value 0 with 2,000 resamples).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Normalised pivotQuantity compared.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
subject value
0.2128 normalised shareMean of the metric for the subject.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
comparator value
0.9838 normalised shareMean of the metric for the comparator.
value
-0.771 normalised shareMean over questions of subject minus comparator.
interval low
-0.7764 normalised shareLower bound of the 95% paired bootstrap interval.
interval high
-0.7655 normalised shareUpper bound of the 95% paired bootstrap interval.
p holm
0 p-valueHolm-adjusted two-sided bootstrap p-value over A1, A2 and B.
scope
766 of 800 Polish STEM questions usable in both models, answers with Polish chain-of-thoughtQuestions and answers compared.

Experimental provenance

Method and evaluation protocol
Teacher-forced forward pass of each answer through the MLX bf16 model with a custom loop over the 50 blocks; after each block the residual stream at generated positions is decoded through the final RMSNorm and output head with an fp32 softmax; masses on English- and Polish-labelled pieces are pooled over reasoning positions whose next token is language-labelled. Answers below 0.95 teacher-forcing agreement, with a failed round-trip or without language positions are excluded. Anchors per model and benchmark from the pooled layer-50 English share of each condition; per question, the transplant's normalised pivot minus the original's.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
Means 0.2128 and 0.9838; difference -0.771 [-0.7764, -0.7655], Holm p 0. Anchors (layer-50 English share): transplant gsm8k_pl pl 0.1096; transplant gsm8k_pl en 0.965; transplant llmzszl_stem pl 0.1095; transplant llmzszl_stem en 0.948; transplant pes pl 0.0232; transplant pes en 0.7852; original gsm8k_pl pl 0.0236; original gsm8k_pl en 0.9218; original llmzszl_stem pl 0.0239; original llmzszl_stem en 0.8274; original pes pl 0.006; original pes en 0.4922. Difference in pivot excess -0.5098 [-0.5174, -0.5027].
Uncertainty and replication
95% stratified paired item bootstrap interval (2000 resamples, strata benchmark, seed 20260704); Holm adjustment over A1, A2 and B. A p-value of 0 means that no bootstrap resample reached zero; with 2000 resamples the smallest non-zero value is 0.001. Sensitivity: truncation-excluded -0.7709 [-0.7763, -0.7653] (741 items); compliance-restricted -0.776 [-0.7815, -0.7703] (733 items); agreement at least 0.98 -0.7341 [-0.7403, -0.7278] (498 items).
Limitations
Logit lens, not a tuned lens: early-layer decodings are noisy. The vocabulary partitions label 16163 Polish and 1289 English pieces for APT4 and 1665 and 6116 for the Mistral-derived tokenizer, and were built from re-serialised MLX tokenizer files without an equivalence check. The design was written after the lens run started. The Polish-model lens file combines records from two invocations separated by an unrecorded rewrite of the lens script. Anchors correct the level bias of the partitions only at layer 50, so the magnitude depends on the partitions. The design registered A2 as descriptive and directional, yet it enters the Holm family.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Normalised pivot

For one answer, the band English share minus the Polish anchor, divided by the English anchor minus the Polish anchor, where the anchors are the model's pooled layer-50 English shares on its answers with Polish and with English chain-of-thought to the questions of the same benchmark; 0 is the model's overt-Polish level and 1 its overt-English level.

Key normalised_pivot · version e142a14d-192a-4ae7-ad70-c2300ef06faf

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

related

fd782265-b2b7-4324-851f-ade0cb5eff31

The pivot excess of the transplant on the same answers.