Immutable accepted plan · retrospective
Reasoning-language forcing: two Bielik 11B models under Polish and English chain-of-thought on 800 Polish STEM questions
Amendment 1, reconstructed in the port: max_tokens for GSM8K-PL raised from 1024 to 1280 before generation (4 July 2026, 00:03 +0200); the original never wrote it into the design.
Public source
https://github.com/stw2/tokenizer-science-tax @ 02380f96e12e849fae306c5e9bc44f12c7e7a866
Reference checked 2026-09-14 13:10 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- The English chain-of-thought gain of Bielik-PL-11B-v3.0-Instruct minus that of Bielik-11B-v3.0-Instruct, pooled over the 800 questions, is above zero; the gain is larger where Polish reasoning traces take more tokens than English ones.
- Protocol
- 01 builds the three item sets from pinned dataset revisions (GSM8K-PL on the 250 test indices of E2's bridge specification after a check that Polish and English gold answers agree; LLMzSzŁ STEM by subject and PES by specialization, proportional stratified sampling) and the few-shot exemplars, with reasoning traces written by hand for the multiple-choice exemplars; 02 runs a smoke gate (BOS parity, byte-identical greedy repeat, throughput projection, language check of 12 outputs), then answers every item under Polish and English chain-of-thought with both models, the question always in Polish; 03 grades the answer after the last final-answer marker (four hash signs), with counted fallback extraction, and classifies the reasoning language; 04 computes accuracies, the English chain-of-thought gain per model with exact McNemar tests, the difference in gains with a stratified paired item bootstrap, per-benchmark differences with Holm correction, compliance-restricted and truncation-excluded sensitivity analyses and generated-token counts.
- Dataset
- GSM8K-PL: 250 problems of the Eurolingua/gsm8kx Polish GSM8K test split with openai/gsm8k gold answers. LLMzSzŁ STEM: 300 questions in mathematics, physics, science and biology from the amu-cai/llmzszl-dataset test split, vocational examinations and questions referring to figures or tables excluded. PES: 250 questions of the Polish State Specialization Examination 2018 to 2022 over 57 specializations.
- Split
- Evaluation only: test-split problems (GSM8K-PL, LLMzSzŁ) and examination questions (PES); few-shot exemplars from the gsm8kx few-shot pool and held out of the multiple-choice pools.
- Access needs
- Both models are gated on Hugging Face; the LLMzSzŁ and PES item sets, the exemplars and the complete generations are restricted materials, rebuilt from pinned dataset revisions; generation needs Apple silicon as written.
- Configurations
- Conditions: system prompt, few-shot reasoning traces and scaffold labels in Polish or in English, question in Polish. Few-shot: 5 exemplars for GSM8K-PL, 3 per multiple-choice benchmark. max_tokens 1280 for GSM8K-PL and 768 for multiple-choice. Greedy decoding, MLX bf16.
- Metric
- Answer accuracy; English chain-of-thought gain per model (English minus Polish accuracy) with a 95% bootstrap interval and exact McNemar test; DD, the gain of Bielik-PL-11B-v3.0-Instruct minus the gain of Bielik-11B-v3.0-Instruct, pooled with a 95% stratified paired item bootstrap interval; mean generated tokens per answer.
- Seeds
- 20260703 for item selection, exemplar draws and the bootstrap (2000 resamples); decoding is greedy.
- Interpretation rule
- Primary: DD pooled over the 800 items, intention-to-treat; H1 confirmed if the 95% interval lies above zero, refuted (inverted) if it lies below zero, otherwise inconclusive. Secondary: per-benchmark DDs with Holm correction over three; per-model gains with exact McNemar tests; compliance-restricted and truncation-excluded reruns of the primary. H2: descriptive ordering of gain against the Polish trace token surplus over six benchmark and model combinations, weak evidence either way, and an exploratory item-level association.
- Resources
- Apple silicon with MLX, one 11B bf16 model at a time; about half a day of generation per model.
- Prior work
- arXiv:2604.10799v1 reports GSM8K 80.97 for Bielik-PL-11B-v3.0-Instruct against 85.60 for Bielik-11B-v3.0-Instruct, and Polish Medical Leaderboard 48.42 against 50.21 percent.
Selected exact hypotheses and premises
Premise · P34
e9fd6b0d-e665-41be-a6dc-290547a16735Polish Medical Leaderboard values of the model pair, built on PES questions.
Premise · P16
8368afe3-5846-42c2-9352-ffa0ff2db353Polish fertility of APT4 and the Mistral-derived tokenizer.