Experiment proposal

Sign in with GitHub
← Current experiment E3

Exact proposal revision

Does forcing English chain-of-thought on Polish STEM questions raise the accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct?

Proposed by @stw2 via agent · 2026-09-14 13:10 UTC

800 questions posed in Polish (250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions, 250 PES examination questions), each answered by both models under Polish and English chain-of-thought with greedy decoding.

Access and suggested protocol

Access needs
Both models are gated on Hugging Face; the LLMzSzŁ and PES item sets, the exemplars and the complete generations are restricted materials, rebuilt from pinned dataset revisions; generation needs Apple silicon as written.
Suggested protocol
Scripts 01 to 08 in experiments/E03-reasoning-language, in order.

Selected exact hypotheses and premises

Hypothesis · H8

0a102e08-95f9-4674-840b-802a26c1f2d6

Primary prediction under test.

Hypothesis · H9

b34419c3-aeeb-4796-b264-a825baa0a9f7

Secondary, mechanism prediction.

Premise · P29

36da68d9-97cf-4575-ad42-ed56e0e895e9

The regression whose mechanism is probed.

Reason for this revision

Initial proposal.