Immutable accepted plan · retrospective
Teacher-forced logit lens of the Bielik 11B pair over existing Polish and English chain-of-thought traces, with anchor-calibrated pivot metrics and three tests
Amendment 1, dated 4 July 2026 and written after the first 69 full-run records, before any test: records below 0.95 teacher-forcing agreement are excluded from the analysis instead of the 0.98 per-record gate, the mean smoke gate stays at 0.98, and a sensitivity arm restricted to records at 0.98 or above is added to A1, A2 and B.
Public source
https://github.com/stw2/tokenizer-science-tax @ 532f491ba531d07f992b45a1465f48e74a8b058a
Reference checked 2026-09-14 13:12 UTC. No code was executed or scientific result verified.
Plan
- Prediction
- For Bielik-PL-11B-v3.0-Instruct on Polish STEM questions with Polish chain-of-thought, the mean pivot excess is above zero, and the band English share has a positive logistic coefficient for answer correctness.
- Protocol
- 00 checks the inputs read from other experiments; 01 builds the English, Polish and other vocabulary partition of each tokenizer from five corpora and checks golden pieces; 02 runs a smoke gate, then one teacher-forced forward pass per trace that decodes every block's residual stream at generated positions through the final normalisation and output head and stores per-layer English and Polish masses and argmax counts over reasoning positions; 03 computes per-item pivot metrics over layers 26 to 43, per-model and per-benchmark final-layer anchors, tests A1, A2 and B with a stratified paired bootstrap and Holm adjustment, and sensitivity arms; 04 writes flat headline values; 05 writes the content manifest.
- Dataset
- 3,200 greedy traces of the two models on 800 Polish items (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) in Polish and English chain-of-thought conditions, with graded outcomes; five corpora for the vocabulary partitions (GSM8K questions and arXiv abstracts in English; Polish Wikipedia science articles, reviews and PES questions).
- Split
- None: the Polish-condition traces of items usable in both models form the test universe; the English-condition traces give the anchors.
- Access needs
- Both models are gated with automatic approval; as written, the lens runs need Apple silicon, 128 GB unified memory; the complete answers, the PES and LLMzSzŁ items, the few-shot exemplars and two corpora are restricted materials.
- Configurations
- Band layers 26 to 43 of 50, sensitivity bands 13 to 25 and 30 to 40; vocabulary partition minimum count 5, rate ratio 10, Polish-diacritic override; gates: BOS parity, round-trip at least 99%, mean teacher-forcing agreement at least 0.98, lens final layer against model logits, determinism on repeated passes. Amendment 1: per-record exclusion below 0.95 teacher-forcing agreement; strict arm at 0.98.
- Metric
- Pivot excess (band English share minus final-layer English share); normalised pivot ((band English share minus the Polish anchor) divided by (English anchor minus Polish anchor)); logistic coefficient of correctness on the standardised band English share.
- Seeds
- 20260704 (bootstrap and audit samples); 2000 bootstrap resamples
- Interpretation rule
- A1: two-sided stratified paired bootstrap of the transplant's mean pivot excess, Holm-adjusted within the family A1, A2, B. A2: paired difference of normalised pivot between the models, descriptive and directional. B: sign and 95% bootstrap interval of the coefficient, with a difficulty-controlled secondary and per-benchmark point-biserial correlations. No decision bins or equivalence bounds are registered. Records below 0.95 teacher-forcing agreement, with a failed round-trip or without language-labelled positions are excluded; A1, A2 and B are repeated on records at 0.98 or above.
- Resources
- Apple silicon, 128 GB unified memory; about two hours of MLX forward passes; no training and no new generations.
- Prior work
- Logit-lens studies of Llama-family models report an English-like latent language in middle layers (arXiv:2402.10588).
Selected exact hypotheses and premises
Premise · P29
36da68d9-97cf-4575-ad42-ed56e0e895e9The reasoning regression whose mechanism is probed.
Premise · P76
a8a88632-e5bd-42b6-8a04-17b77ce87d13The chain-of-thought language test on the same answers, whose surface mechanism this probe follows inward.