Experiment · E8
Does Bielik-PL-11B-v3.0-Instruct still pass through an English-like latent space when it reasons in Polish on Polish mathematics and STEM problems, and does the strength of that pivot predict whether its answer is correct?
Teacher-forced logit lens of Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct over 3,200 existing greedy traces: 800 Polish items (GSM8K-PL, LLMzSzŁ STEM, PES) in Polish and English chain-of-thought conditions; per-tokenizer vocabulary language partitions from five corpora; no new generations.
Prerequisites and protocol
- Access needs
- Both models are gated with automatic approval; as written, the lens runs need Apple silicon, 128 GB unified memory; the complete answers, the PES and LLMzSzŁ items, the few-shot exemplars and two corpora are restricted materials.
- Suggested protocol
- Scripts 00 to 05 in experiments/E08-latent-pivot, in order.
Accepted plan
Plan accepted by @stw2
Teacher-forced logit lens of the Bielik 11B pair over existing Polish and English chain-of-thought traces, with anchor-calibrated pivot metrics and three tests- requires model repository bielik-pl-11b-v3.0-instruct · model · M46 Bielik-PL-11B-v3.0-Instruct · Ask the reporter
- requires model repository bielik-11b-v3.0-instruct · model · M47 Bielik-11B-v3.0-Instruct · Ask the reporter
- requires bielik-pl-11b-v3.0-instruct mlx-bf16 · checkpoint · M65 bielik-pl-11b-v3.0-instruct mlx-bf16 · Ask the reporter
- requires bielik-11b-v3.0-instruct mlx-bf16 · checkpoint · M66 bielik-11b-v3.0-instruct mlx-bf16 · Ask the reporter
- requires MLX tokenizer.json bielik-pl-11b-v3.0-instruct · tokenizer · M394 MLX tokenizer.json of Bielik-PL-11B-v3.0-Instruct · Ask the reporter
- requires MLX tokenizer.json bielik-11b-v3.0-instruct · tokenizer · M395 MLX tokenizer.json of Bielik-11B-v3.0-Instruct · Ask the reporter
- requires corpus en-gsm8k · reference corpus · M30 en-gsm8k.jsonl · Download
- requires corpus en-arxiv-abstracts · reference corpus · M27 en-arxiv-abstracts.jsonl · Download
- requires corpus pl-wiki-science · reference corpus · M32 pl-wiki-science.jsonl · Download
- requires corpus pl-informal · reference corpus · M33 pl-informal.jsonl · Ask the reporter
- requires corpus pl-science-pes · reference corpus · M31 pl-science-pes.jsonl · Ask the reporter
- requires traces bielik-pl-11b-v3.0-instruct · evaluation data · M99 raw/run_bielik-pl-11b-v3.0-instruct.jsonl · Ask the reporter
- requires traces bielik-11b-v3.0-instruct · evaluation data · M100 raw/run_bielik-11b-v3.0-instruct.jsonl · Ask the reporter
- requires graded outcomes bielik-pl-11b-v3.0-instruct · evaluation data · M102 graded_bielik-pl-11b-v3.0-instruct.jsonl · Download
- requires graded outcomes bielik-11b-v3.0-instruct · evaluation data · M103 graded_bielik-11b-v3.0-instruct.jsonl · Download
- requires prompt builder 02_run.py · prompt builder · M396 Prompt builder of the reasoning-language runs, version run with the lens · No known location
- requires e3_config.json · configuration · M397 Configuration of the reasoning-language runs · Download
- requires item manifest · configuration · M96 manifest.json · Download
- requires exemplars · evaluation data · M95 exemplars.json · Ask the reporter
- requires items gsm8k-pl · evaluation data · M92 items_gsm8k_pl.jsonl · Download
- requires items llmzszl-stem · evaluation data · M93 items_llmzszl_stem.jsonl · Ask the reporter
- requires items pes · evaluation data · M94 items_pes.jsonl · Ask the reporter
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 6
6 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 6a20c5c41c50 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 532f491ba531 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ a9246a60690b · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →