Accepted plan

Sign in with GitHub
← Experiment E3

Immutable accepted plan · retrospective

Reasoning-language forcing: two Bielik 11B models under Polish and English chain-of-thought on 800 Polish STEM questions

Accepted by @stw2 via agent. Planning intent is the author’s declaration.

Amendment 1, reconstructed in the port: max_tokens for GSM8K-PL raised from 1024 to 1280 before generation (4 July 2026, 00:03 +0200); the original never wrote it into the design.

Public source

Plan

Prediction
The English chain-of-thought gain of Bielik-PL-11B-v3.0-Instruct minus that of Bielik-11B-v3.0-Instruct, pooled over the 800 questions, is above zero; the gain is larger where Polish reasoning traces take more tokens than English ones.
Protocol
01 builds the three item sets from pinned dataset revisions (GSM8K-PL on the 250 test indices of E2's bridge specification after a check that Polish and English gold answers agree; LLMzSzŁ STEM by subject and PES by specialization, proportional stratified sampling) and the few-shot exemplars, with reasoning traces written by hand for the multiple-choice exemplars; 02 runs a smoke gate (BOS parity, byte-identical greedy repeat, throughput projection, language check of 12 outputs), then answers every item under Polish and English chain-of-thought with both models, the question always in Polish; 03 grades the answer after the last final-answer marker (four hash signs), with counted fallback extraction, and classifies the reasoning language; 04 computes accuracies, the English chain-of-thought gain per model with exact McNemar tests, the difference in gains with a stratified paired item bootstrap, per-benchmark differences with Holm correction, compliance-restricted and truncation-excluded sensitivity analyses and generated-token counts.
Dataset
GSM8K-PL: 250 problems of the Eurolingua/gsm8kx Polish GSM8K test split with openai/gsm8k gold answers. LLMzSzŁ STEM: 300 questions in mathematics, physics, science and biology from the amu-cai/llmzszl-dataset test split, vocational examinations and questions referring to figures or tables excluded. PES: 250 questions of the Polish State Specialization Examination 2018 to 2022 over 57 specializations.
Split
Evaluation only: test-split problems (GSM8K-PL, LLMzSzŁ) and examination questions (PES); few-shot exemplars from the gsm8kx few-shot pool and held out of the multiple-choice pools.
Access needs
Both models are gated on Hugging Face; the LLMzSzŁ and PES item sets, the exemplars and the complete generations are restricted materials, rebuilt from pinned dataset revisions; generation needs Apple silicon as written.
Configurations
Conditions: system prompt, few-shot reasoning traces and scaffold labels in Polish or in English, question in Polish. Few-shot: 5 exemplars for GSM8K-PL, 3 per multiple-choice benchmark. max_tokens 1280 for GSM8K-PL and 768 for multiple-choice. Greedy decoding, MLX bf16.
Metric
Answer accuracy; English chain-of-thought gain per model (English minus Polish accuracy) with a 95% bootstrap interval and exact McNemar test; DD, the gain of Bielik-PL-11B-v3.0-Instruct minus the gain of Bielik-11B-v3.0-Instruct, pooled with a 95% stratified paired item bootstrap interval; mean generated tokens per answer.
Seeds
20260703 for item selection, exemplar draws and the bootstrap (2000 resamples); decoding is greedy.
Interpretation rule
Primary: DD pooled over the 800 items, intention-to-treat; H1 confirmed if the 95% interval lies above zero, refuted (inverted) if it lies below zero, otherwise inconclusive. Secondary: per-benchmark DDs with Holm correction over three; per-model gains with exact McNemar tests; compliance-restricted and truncation-excluded reruns of the primary. H2: descriptive ordering of gain against the Polish trace token surplus over six benchmark and model combinations, weak evidence either way, and an exploratory item-level association.
Resources
Apple silicon with MLX, one 11B bf16 model at a time; about half a day of generation per model.
Prior work
arXiv:2604.10799v1 reports GSM8K 80.97 for Bielik-PL-11B-v3.0-Instruct against 85.60 for Bielik-11B-v3.0-Instruct, and Polish Medical Leaderboard 48.42 against 50.21 percent.

Selected exact hypotheses and premises

Hypothesis · H8

0a102e08-95f9-4674-840b-802a26c1f2d6

Primary prediction under test.

Hypothesis · H9

b34419c3-aeeb-4796-b264-a825baa0a9f7

Secondary, mechanism prediction.

Premise · P29

36da68d9-97cf-4575-ad42-ed56e0e895e9

GSM8K values of the model pair.

Premise · P34

e9fd6b0d-e665-41be-a6dc-290547a16735

Polish Medical Leaderboard values of the model pair, built on PES questions.

Premise · P16

8368afe3-5846-42c2-9352-ffa0ff2db353

Polish fertility of APT4 and the Mistral-derived tokenizer.