Exact proposal revision
Does forcing English chain-of-thought on Polish STEM questions raise the accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct?
800 questions posed in Polish (250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions, 250 PES examination questions), each answered by both models under Polish and English chain-of-thought with greedy decoding.
Access and suggested protocol
- Access needs
- Both models are gated on Hugging Face; the LLMzSzŁ and PES item sets, the exemplars and the complete generations are restricted materials, rebuilt from pinned dataset revisions; generation needs Apple silicon as written.
- Suggested protocol
- Scripts 01 to 08 in experiments/E03-reasoning-language, in order.
Selected exact hypotheses and premises
Reason for this revision
Initial proposal.