Finding

Sign in with GitHub
← Publications

Finding · P201 · Author-curated

On 766 of 800 Polish STEM questions with Polish chain-of-thought, the mean band argmax English share of Bielik-PL-11B-v3.0-Instruct minus that of Bielik-11B-v3.0-Instruct is -0.0501 (95% interval -0.0564 to -0.0436).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Paired difference

metric
Band argmax English shareQuantity compared.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
value
-0.0501 shareMean over questions of subject minus comparator.
interval low
-0.0564 shareLower bound of the 95% paired bootstrap interval.
interval high
-0.0436 shareUpper bound of the 95% paired bootstrap interval.
scope
766 of 800 Polish STEM questions usable in both models, answers with Polish chain-of-thoughtQuestions and answers compared.

Experimental provenance

Method and evaluation protocol
Teacher-forced forward pass of each answer through the MLX bf16 model with a custom loop over the 50 blocks; after each block the residual stream at generated positions is decoded through the final RMSNorm and output head with an fp32 softmax; masses on English- and Polish-labelled pieces are pooled over reasoning positions whose next token is language-labelled. Answers below 0.95 teacher-forcing agreement, with a failed round-trip or without language positions are excluded. Per layer, argmax token labels instead of probability masses; per question, band mean; paired difference between the models.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
-0.0501 [-0.0564, -0.0436].
Uncertainty and replication
95% stratified paired item bootstrap interval (2000 resamples, strata benchmark, seed 20260704); a sensitivity variant, not in the Holm family.
Limitations
Logit lens, not a tuned lens: early-layer decodings are noisy. The vocabulary partitions label 16163 Polish and 1289 English pieces for APT4 and 1665 and 6116 for the Mistral-derived tokenizer, and were built from re-serialised MLX tokenizer files without an equivalence check. The design was written after the lens run started. The Polish-model lens file combines records from two invocations separated by an unrecorded rewrite of the lens script. Added in the analysis script; not registered in the design.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Band argmax English share

For one answer, averaged over layers 26 to 43 of a 50-layer model: among the positions whose next token is labelled English or Polish, the number whose highest-probability logit-lens token is labelled English divided by the number whose highest-probability token is labelled English or Polish.

Key band_argmax_english_share · version 596d7e01-9a33-4e02-9e74-4139d5d42c9a

Paired difference

value is the subject's metric minus the comparator's metric over the same items in scope, averaged over settings where listed; interval_low and interval_high bound its 95% paired bootstrap interval; subject_value and comparator_value are the two metrics, benchmark names the benchmark and p_holm the Holm-adjusted p-value, where given.

Key paired_difference · version d7c031db-103f-410d-b439-1f3ba6cae904

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

supports

e142a14d-192a-4ae7-ad70-c2300ef06faf

Same direction with hard decoding, which does not depend on diffuse probability mass.