Finding

Sign in with GitHub
← Publications

Finding · P202 · Author-curated

On Polish STEM questions, the mean band English share of Bielik-PL-11B-v3.0-Instruct is 0.258 for its answers with Polish chain-of-thought and 0.3671 for its answers with English chain-of-thought.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric comparison

metric
Band English shareQuantity compared.
subject
Polish chain-of-thoughtPrompting condition of the first value.
subject value
0.258 shareMean band English share under the subject condition.
comparator
English chain-of-thoughtPrompting condition of the second value.
comparator value
0.3671 shareMean band English share under the comparator condition.
setting
Bielik-PL-11B-v3.0-InstructModel both values are measured on.
scope
Polish chain-of-thought: 766 of 800 Polish STEM questions usable in both models; English chain-of-thought: 767 usable answers of this modelAnswers averaged.

Experimental provenance

Method and evaluation protocol
Teacher-forced forward pass of each answer through the MLX bf16 model with a custom loop over the 50 blocks; after each block the residual stream at generated positions is decoded through the final RMSNorm and output head with an fp32 softmax; masses on English- and Polish-labelled pieces are pooled over reasoning positions whose next token is language-labelled. Answers below 0.95 teacher-forcing agreement, with a failed round-trip or without language positions are excluded. Band English share per answer, averaged per condition.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
0.258 against 0.3671; the report reads the ratio as 70%.
Uncertainty and replication
Point values without intervals.
Limitations
Logit lens, not a tuned lens: early-layer decodings are noisy. The vocabulary partitions label 16163 Polish and 1289 English pieces for APT4 and 1665 and 6116 for the Mistral-derived tokenizer, and were built from re-serialised MLX tokenizer files without an equivalence check. The design was written after the lens run started. The Polish-model lens file combines records from two invocations separated by an unrecorded rewrite of the lens script. The two means average different sets of answers.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

English chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in English; the question is in Polish.

Key english_chain_of_thought · version 5907d582-517e-4a59-bdc6-c58cf7a7fe7d

Band English share

The logit-lens English share averaged over layers 26 to 43 of a 50-layer model, for one answer.

Key band_english_share · version 25a042fa-58a1-4b0a-8984-73e43685a6e0

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Polish chain-of-thought

Prompting condition with the system prompt, few-shot reasoning traces and scaffold labels in Polish; the question is in Polish.

Key polish_chain_of_thought · version d0b79462-c553-4a1f-9612-98ca6d4a0645

Exact references

related

e142a14d-192a-4ae7-ad70-c2300ef06faf

A within-model contrast with the same vocabulary partition.

related

af0896b4-cbbe-43b0-996f-f873cf2ebb93

How often the answers are written in the prompted language under each condition.