Experiment · E3
Does forcing English chain-of-thought on Polish STEM questions raise the accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct?
800 questions posed in Polish (250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions, 250 PES examination questions), each answered by both models under Polish and English chain-of-thought with greedy decoding.
Prerequisites and protocol
- Access needs
- Both models are gated on Hugging Face; the LLMzSzŁ and PES item sets, the exemplars and the complete generations are restricted materials, rebuilt from pinned dataset revisions; generation needs Apple silicon as written.
- Suggested protocol
- Scripts 01 to 08 in experiments/E03-reasoning-language, in order.
Accepted plan
Plan accepted by @stw2
Reasoning-language forcing: two Bielik 11B models under Polish and English chain-of-thought on 800 Polish STEM questions- requires Bielik-PL-11B-v3.0-Instruct · model · M46 Bielik-PL-11B-v3.0-Instruct · Ask the reporter
- requires Bielik-11B-v3.0-Instruct · model · M47 Bielik-11B-v3.0-Instruct · Ask the reporter
- requires gsm8kx Polish test split · item source · M87 gsm8kx Polish GSM8K test split · Download
- requires gsm8kx Polish few-shot pool · item source · M88 gsm8kx Polish GSM8K few-shot pool · Download
- requires GSM8K main configuration · item source · M89 GSM8K main configuration · Download
- requires LLMzSzŁ test split · item source · M90 LLMzSzŁ test split · Download
- requires PES 2018-2022 questions · item source · M16 PES 2018-2022 questions · Download
- requires GSM8K bridge specification · item selection · M78 gsm8k_bridge_spec.json · Download
- requires authored exemplar traces · few shot traces · M91 Authored exemplar reasoning traces · Ask the reporter
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 8
8 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
planned attempt · stw2/tokenizer-science-tax @ c10ec1755e97 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 02380f96e12e · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan
A20 Grades every answer after a self-test gate and classifies the language of its reasoning.
Succeededplanned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan
A23 Copies the header and GSM8K-PL records of the complete generations, byte for byte, to the redistributable files.
Succeededplanned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan
A24 Writes flat headline values from the analysis, the token-cap audit and the item manifest.
Succeededplanned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →