Immutable accepted plan · retrospective
Teacher-forced logit lens of the Bielik 11B pair over existing Polish and English chain-of-thought traces, with anchor-calibrated pivot metrics and three tests
The design as written before its Amendment 1, from DESIGN.md of the archived commit of 4 July 2026, which holds the design and the results together.
Public source
https://github.com/stw2/tokenizer-science-tax @ 6a20c5c41c509137b0de80a80ec4fd3d3243ba56
Reference checked 2026-09-14 13:12 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- For Bielik-PL-11B-v3.0-Instruct on Polish STEM questions with Polish chain-of-thought, the mean pivot excess is above zero, and the band English share has a positive logistic coefficient for answer correctness.
- Protocol
- 00 checks the inputs read from other experiments; 01 builds the English, Polish and other vocabulary partition of each tokenizer from five corpora and checks golden pieces; 02 runs a smoke gate, then one teacher-forced forward pass per trace that decodes every block's residual stream at generated positions through the final normalisation and output head and stores per-layer English and Polish masses and argmax counts over reasoning positions; 03 computes per-item pivot metrics over layers 26 to 43, per-model and per-benchmark final-layer anchors, tests A1, A2 and B with a stratified paired bootstrap and Holm adjustment, and sensitivity arms; 04 writes flat headline values; 05 writes the content manifest.
- Dataset
- 3,200 greedy traces of the two models on 800 Polish items (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) in Polish and English chain-of-thought conditions, with graded outcomes; five corpora for the vocabulary partitions (GSM8K questions and arXiv abstracts in English; Polish Wikipedia science articles, reviews and PES questions).
- Split
- None: the Polish-condition traces of items usable in both models form the test universe; the English-condition traces give the anchors.
- Access needs
- Both models are gated with automatic approval; as written, the lens runs need Apple silicon, 128 GB unified memory; the complete answers, the PES and LLMzSzŁ items, the few-shot exemplars and two corpora are restricted materials.
- Configurations
- Band layers 26 to 43 of 50, sensitivity bands 13 to 25 and 30 to 40; vocabulary partition minimum count 5, rate ratio 10, Polish-diacritic override; gates: BOS parity, round-trip at least 99%, mean teacher-forcing agreement at least 0.98, lens final layer against model logits, determinism on repeated passes.
- Metric
- Pivot excess (band English share minus final-layer English share); normalised pivot ((band English share minus the Polish anchor) divided by (English anchor minus Polish anchor)); logistic coefficient of correctness on the standardised band English share.
- Seeds
- 20260704 (bootstrap and audit samples); 2000 bootstrap resamples
- Interpretation rule
- A1: two-sided stratified paired bootstrap of the transplant's mean pivot excess, Holm-adjusted within the family A1, A2, B. A2: paired difference of normalised pivot between the models, descriptive and directional. B: sign and 95% bootstrap interval of the coefficient, with a difficulty-controlled secondary and per-benchmark point-biserial correlations. No decision bins or equivalence bounds are registered.
- Resources
- Apple silicon, 128 GB unified memory; about two hours of MLX forward passes; no training and no new generations.
- Prior work
- Logit-lens studies of Llama-family models report an English-like latent language in middle layers (arXiv:2402.10588).
Selected exact hypotheses and premises
Premise · P29
36da68d9-97cf-4575-ad42-ed56e0e895e9The reasoning regression whose mechanism is probed.
Premise · P76
a8a88632-e5bd-42b6-8a04-17b77ce87d13The chain-of-thought language test on the same answers, whose surface mechanism this probe follows inward.