Experiment

Sign in with GitHub
← Experiments

Experiment · E3

Does forcing English chain-of-thought on Polish STEM questions raise the accuracy of Bielik-PL-11B-v3.0-Instruct more than that of Bielik-11B-v3.0-Instruct?

Completed · Proposed by @stw2 · Assigned to @stw2

800 questions posed in Polish (250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions, 250 PES examination questions), each answered by both models under Polish and English chain-of-thought with greedy decoding.

Prerequisites and protocol

Access needs
Both models are gated on Hugging Face; the LLMzSzŁ and PES item sets, the exemplars and the complete generations are restricted materials, rebuilt from pinned dataset revisions; generation needs Apple silicon as written.
Suggested protocol
Scripts 01 to 08 in experiments/E03-reasoning-language, in order.

Accepted plan

Plan accepted by @stw2

Reasoning-language forcing: two Bielik 11B models under Polish and English chain-of-thought on 800 Polish STEM questions

retrospective

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts · 8

8 succeeded

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

  1. planned attempt · stw2/tokenizer-science-tax @ c10ec1755e97 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  2. planned attempt · stw2/tokenizer-science-tax @ 02380f96e12e · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  3. planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  4. planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  5. planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  6. planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  7. planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  8. planned attempt · stw2/tokenizer-science-tax @ d3c436d89b85 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →