Public research
Publications
Cited claims and findings from everyone in The tokenizer science tax. Each row opens the exact published assertion and its provenance.
Wording
Reanalysis of the released scores for 800 Polish documents from the 1.5B APT4 recovery-training confirmation finds a 30%-versus-0% mathematics/code cost of 0.015635 bits per byte. A paired document bootstrap gives a 95% interval of [0.011866, 0.018028], entirely below the 0.05 bits-per-byte bound.31354266-8c39-4d95-bcd9-8738784d3fdc · author-curated
Executing the pinned E6 analysis programs on their released score arrays and digit-probe outcomes under macOS arm64, Python 3.12.13 and NumPy 1.26.3 reproduces all 101 result fields of analysis.json and analysis_conf.json, with zero absolute numeric difference.87190d44-afa8-402d-82e3-045f5a27d523 · author-curated
Training the Polish-only 32k tokenizer a second time from the same input gives identical model content once the stored output path is cleared and a byte-identical tokenizer.json; the raw model files differ.a84fdc29-bf64-4a26-8ac6-6a5f5d029008 · author-curated
All 5 arms pass the structural, digit-policy, fidelity and separator-piece gates, with 0 piece and 0 decode mismatches between SentencePiece and the fast tokenizer on 10,000 lines per arm.f1aa77fe-5c2d-441e-a2e3-a68cdcbeb49d · author-curated
Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer after 500M tokens of continued pretraining on the Polish questions of 600 MMLU translation pairs from 0.3617 to 0.4200, a gain of 0.0583 (95% interval 0.0266 to 0.0917) that recovers 0.6250 of its English-minus-Polish accuracy gap of 0.0933.f809bfdc-15e1-4b59-9fe1-89642837e43f · author-curated
Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer after 500M tokens of continued pretraining on the Polish questions of 600 Belebele translation pairs from 0.4767 to 0.5300, a gain of 0.0533 (95% interval 0.0133 to 0.0917) that recovers 0.3951 of its English-minus-Polish accuracy gap of 0.1350.72ec1fd8-81ee-49b3-9f77-f464d8e50ec3 · author-curated
Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer on the Polish questions of 600 MMLU translation pairs from 0.3417 to 0.4467, a gain of 0.1050 (95% interval 0.0650 to 0.1467) that recovers 0.9692 of its English-minus-Polish accuracy gap of 0.1083.2a4f4dc0-fc2a-422c-9ffc-13ad60ecabcf · author-curated
Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer on the Polish questions of 600 Belebele translation pairs from 0.4300 to 0.5933, a gain of 0.1633 (95% interval 0.1217 to 0.2050) that recovers 0.9074 of its English-minus-Polish accuracy gap of 0.1800.54ed129a-1b1a-499a-871e-dc4537e4d95b · author-curated
Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B on the Polish questions of 600 MMLU translation pairs from 0.3400 to 0.5100, a gain of 0.1700 (95% interval 0.1250 to 0.2150) that recovers 0.7969 of its English-minus-Polish accuracy gap of 0.2133.b1699637-dcb7-40e3-836e-4de6c3d4eb01 · author-curated
Prepending the English twin's prompt raises the likelihood-scored accuracy of Qwen2.5-1.5B on the Polish questions of 600 Belebele translation pairs from 0.5483 to 0.6367, a gain of 0.0883 (95% interval 0.0483 to 0.1317) that recovers 0.4454 of its English-minus-Polish accuracy gap of 0.1983.803ee4c6-d174-4db8-be9d-f81f6313c7f1 · author-curated
Swapping the answer-scaffold language changes the likelihood-scored accuracy of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer after 500M tokens of continued pretraining by 0.0250 on Polish and -0.0483 on English questions of 600 Belebele translation pairs, and by 0.0383 and -0.0600 on 600 MMLU translation pairs; not every absolute change is below 0.03.1e57e8aa-363d-43c4-acea-12254d3a00d3 · author-curated
Swapping the answer-scaffold language changes the likelihood-scored accuracy of Qwen2.5-1.5B after 500M tokens of continued pretraining with its own tokenizer by 0.0167 on Polish and 0.0017 on English questions of 600 Belebele translation pairs, and by 0.0183 and -0.0233 on 600 MMLU translation pairs; every absolute change is below 0.03.a70b9be6-37fd-4862-b8d0-7077f2b5a2c7 · author-curated
Swapping the answer-scaffold language changes the likelihood-scored accuracy of Qwen2.5-1.5B by 0.0233 on Polish and -0.0017 on English questions of 600 Belebele translation pairs, and by 0.0150 and 0.0033 on 600 MMLU translation pairs; every absolute change is below 0.03.6d5dcdcb-65c6-4efe-a0f8-67c7a935084b · author-curated
At source layer 1 and target layer 27 of Qwen2.5-1.5B, the cross-lingual residual patch has flip rate 0.3098 at the scaffold-final token, 0.0399 at the question-final token, 0.0336 at the entity span, 0.0337 at a random position and 0.0368 at the first token, on 326 forward-discordant pairs (149 with an entity span).167c3527-dba7-4c3c-8307-886b508c7934 · author-curated
At source layer 25 and target layer 25 of Qwen2.5-1.5B, the cross-lingual residual patch has flip rate 0.9663 at the scaffold-final token, 0.0153 at the question-final token, 0.0268 at the entity span, 0.0245 at a random position and 0.0184 at the first token, on 326 forward-discordant pairs (149 with an entity span).29322188-0524-4842-b47e-bebde7b39ff6 · author-curated
The scaffold-final residual patch on Qwen2.5-1.5B fails the pre-registered validity gate at all 3 layer pairs confirmed on 326 forward-discordant pairs, with breakage of 0.1600 at 25 into 25, 0.1600 at 25 into 23, 0.7400 at 1 into 27 against at most 0.05, so the patch does not classify the English-minus-Polish accuracy gap as routing-bound or representation-bound.f099f3b9-167a-40e1-926c-023c37b5639d · author-curated
For the scaffold-final residual patch on 100 forward-discordant pairs of Qwen2.5-1.5B, 0 of the 152 odd-layer source and target pairs, among 196, whose transport rate is at most 0.35 have a flip rate at least 3 times their irrelevant-source flip rate; the highest ratio is 1.7500, the highest flip rate 0.3800, and the mean flip rate 0.1747 against a mean irrelevant-source flip rate of 0.1661.fd3ce984-e154-4837-9fbc-642cc7d1024b · author-curated
For the scaffold-final residual patch on 100 forward-discordant pairs of Qwen2.5-1.5B, 35 of the 44 odd-layer source and target pairs, among 196, whose transport rate is above 0.35 have a flip rate at least 3 times their irrelevant-source flip rate; the highest ratio is 5.8750, the highest flip rate 0.9800, and the mean flip rate 0.7557 against a mean irrelevant-source flip rate of 0.1784.3c77469f-74ee-409f-9083-ef12d768623c · author-curated
At source layer 1 and target layer 27 of Qwen2.5-1.5B, the scaffold-final residual patch has flip rate 0.3098, irrelevant-source flip rate 0.3313 and transport rate 0.3098 on 326 forward-discordant pairs, and breakage 0.7400 on 100 Polish-correct pairs; the transport rate is at most 0.35.fbf82ca3-cdab-4e89-8c08-cd37a71f9965 · author-curated
At source layer 25 and target layer 23 of Qwen2.5-1.5B, the scaffold-final residual patch has flip rate 0.9509, irrelevant-source flip rate 0.2393 and transport rate 0.9540 on 326 forward-discordant pairs, and breakage 0.1600 on 100 Polish-correct pairs; the transport rate is above 0.35.71be505b-4844-4c4e-8313-688e1d3d6c3c · author-curated
At source layer 25 and target layer 25 of Qwen2.5-1.5B, the scaffold-final residual patch has flip rate 0.9663, irrelevant-source flip rate 0.2362 and transport rate 0.9632 on 326 forward-discordant pairs, and breakage 0.1600 on 100 Polish-correct pairs; the transport rate is above 0.35.b4e746ee-129b-4085-b633-1c90076cfaa2 · author-curated
On 589 translation-paired MMLU items, Qwen2.5-1.5B's English likelihood multiple-choice accuracy is 0.3964 on the 111 STEM pairs and 0.5941 on the 478 non-STEM pairs.e22c87a2-8e1c-4cb0-b8cf-147b19759f59 · author-curated
In an exploratory split of 589 translation-paired MMLU items, the English-minus-Polish accuracy gap of Qwen2.5-1.5B with APT4 by FVT after 500M tokens of Polish-heavy continued pretraining is 0.0721 (95% interval -0.0057 to 0.1499) on the 111 STEM pairs and 0.1004 (0.0596 to 0.1413) on the 478 non-STEM pairs, a difference of -0.0283.817534c1-b0c7-4c3e-9b82-b4e42375041a · author-curated
In an exploratory split of 589 translation-paired MMLU items, the English-minus-Polish accuracy gap of Qwen2.5-1.5B after 500M tokens of Polish-heavy continued pretraining with its original tokenizer is 0.0721 (95% interval -0.0016 to 0.1458) on the 111 STEM pairs and 0.1192 (0.0721 to 0.1664) on the 478 non-STEM pairs, a difference of -0.0472.af57c29c-3ae6-4934-ab04-3aba003e50fc · author-curated
In an exploratory split of 589 translation-paired MMLU items, the English-minus-Polish accuracy gap of Qwen2.5-1.5B is 0.0541 (95% interval -0.0453 to 0.1534) on the 111 STEM pairs and 0.2510 (0.1989 to 0.3032) on the 478 non-STEM pairs, a difference of -0.1970.7960d8ed-895f-4679-82fd-033f5e8e77f6 · author-curated
After the same 500M tokens of Polish-heavy continued pretraining, the English-minus-Polish accuracy gap of Qwen2.5-1.5B with APT4 by FVT is wider than with its original tokenizer on neither translation-paired Belebele items (difference -0.0418, 95% interval -0.0903 to 0.0100) nor translation-paired MMLU items (difference -0.0153, 95% interval -0.0662 to 0.0357).ceb59c05-b6da-450f-912b-6761f32b457c · author-curated
On 589 translation-paired MMLU items, continued pretraining of Qwen2.5-1.5B on 500M Polish-heavy tokens with its original tokenizer narrows its English-minus-Polish accuracy gap by 0.1036 while English accuracy changes by -0.1036 (95% interval -0.1409 to -0.0645) and Polish accuracy by 0.0000 (-0.0458 to 0.0424): erosion, not improved Polish access.20eefc46-5718-45f8-bf30-405ada15f1c8 · author-curated
After 500M tokens of Polish-heavy continued pretraining with its original tokenizer, Qwen2.5-1.5B's Polish likelihood multiple-choice accuracy rises on neither translation-paired Belebele items (difference -0.1171, 95% interval -0.1605 to -0.0702) nor translation-paired MMLU items (difference 0.0000, 95% interval -0.0458 to 0.0424).72f15189-34bd-468a-a7dd-44743ac5531b · author-curated
On 589 translation-paired MMLU items, after the same 500M tokens of Polish-heavy continued pretraining, the English likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.4584 with APT4 by FVT and 0.4533 with its original tokenizer, a difference of 0.0051 (95% interval -0.0374 to 0.0492).2623f0ec-54c1-4c86-a628-0d55aa1f6e32 · author-curated
On 589 translation-paired MMLU items, after the same 500M tokens of Polish-heavy continued pretraining, the Polish likelihood multiple-choice accuracy of Qwen2.5-1.5B is 0.3633 with APT4 by FVT and 0.3430 with its original tokenizer, a difference of 0.0204 (95% interval -0.0357 to 0.0730).a758ebd5-ec2d-45f4-a58d-9ddb028e511b · author-curated
30 loaded