Finding

Sign in with GitHub
← Publications

Finding · P209 · Author-curated

Averaged over the answers with Polish chain-of-thought to 800 Polish STEM questions, the logit-lens English share of Bielik-11B-v3.0-Instruct is 0.8517 at layer 1, between 0.6275 and 0.7828 over layers 26 to 43, 0.746 at layer 43, 0.0478 at layer 48 and 0.0176 at layer 50.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Layer profile

metric
Logit-lens English shareQuantity profiled.
subject
Bielik-11B-v3.0-InstructModel measured.
evaluation items
Polish STEM question setQuestion set.
condition
Polish chain-of-thoughtPrompting condition of the answers.
items
800 questionsQuestions used.
layer 1
0.8517 shareMean at layer 1.
range layers
26 to 43Layers the range covers.
range min
0.6275 shareLowest mean over range_layers.
range max
0.7828 shareHighest mean over range_layers.
layer 43
0.746 shareMean at layer 43.
layer 48
0.0478 shareMean at layer 48.
layer 50
0.0176 shareMean at layer 50.

Experimental provenance

Method and evaluation protocol
Teacher-forced forward pass of each answer through the MLX bf16 model with a custom loop over the 50 blocks; after each block the residual stream at generated positions is decoded through the final RMSNorm and output head with an fp32 softmax; masses on English- and Polish-labelled pieces are pooled over reasoning positions whose next token is language-labelled. Answers below 0.95 teacher-forcing agreement, with a failed round-trip or without language positions are excluded. Per-layer English share per answer, averaged over the answers.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
Layer 1 0.8517; layers 26 to 43 0.6275 to 0.7828; layer 43 0.746; layer 48 0.0478; layer 50 0.0176.
Uncertainty and replication
Means without intervals.
Limitations
Early-layer logit-lens decodings are noisy. The original's profile averages all its usable Polish-condition answers, a different question set from the transplant's profile.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Layer profile

Averaged over the items used of the evaluation_items, for the subject under the condition, the metric takes the value of each role named layer_<n> at layer n, lies between range_min and range_max over range_layers, and peaks at peak_layer with peak_value, where those roles are given.

Key layer_profile · version dec759b1-6596-420d-ace2-e488f17be61a

Polish chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

Key polish_chain_of_thought · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Polish STEM question set

800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.

Key polish_stem_questions · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Logit-lens English share

At one layer of a model, for the positions of an answer fed back to the model with teacher forcing whose next token is labelled English or Polish by the vocabulary language partition, the probability mass that the model's final normalisation and output head assign to English-labelled pieces when applied to that layer's residual stream, divided by the mass on English- or Polish-labelled pieces, pooled over the positions before the final-answer marker.

Key logit_lens_english_share · version 25a042fa-58a1-4b0a-8984-73e43685a6e0

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c