Experiment

Sign in with GitHub
← Experiments

Experiment · E8

Does Bielik-PL-11B-v3.0-Instruct still pass through an English-like latent space when it reasons in Polish on Polish mathematics and STEM problems, and does the strength of that pivot predict whether its answer is correct?

Completed · Proposed by @stw2 · Assigned to @stw2

Teacher-forced logit lens of Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct over 3,200 existing greedy traces: 800 Polish items (GSM8K-PL, LLMzSzŁ STEM, PES) in Polish and English chain-of-thought conditions; per-tokenizer vocabulary language partitions from five corpora; no new generations.

Prerequisites and protocol

Access needs
Both models are gated with automatic approval; as written, the lens runs need Apple silicon, 128 GB unified memory; the complete answers, the PES and LLMzSzŁ items, the few-shot exemplars and two corpora are restricted materials.
Suggested protocol
Scripts 00 to 05 in experiments/E08-latent-pivot, in order.

Accepted plan

Plan accepted by @stw2

Teacher-forced logit lens of the Bielik 11B pair over existing Polish and English chain-of-thought traces, with anchor-calibrated pivot metrics and three tests

retrospective

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts · 6

6 succeeded

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

  1. planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  2. planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  3. planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  4. planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  5. planned attempt · stw2/tokenizer-science-tax @ 532f491ba531 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  6. planned attempt · stw2/tokenizer-science-tax @ a9246a60690b · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →