Exact proposal revision
Does APT4's tokenization of digits and number separators differ from the Mistral-derived tokenizer's, and does the difference account for the lower GSM8K score of Bielik-PL-11B-v3.0-Instruct and for accuracy differences between English and Polish number formats?
The nine tokenizers of Table 1 of arXiv:2604.10799v1 on synthetic number sets; Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct on 2,500 answer-only arithmetic items in four number formats and two instruction languages, and on 250 GSM8K test problems.
Access and suggested protocol
- Access needs
- The two 11B model repositories are gated on Hugging Face (accept the terms, then use a token); the downloaded snapshots and MLX conversions are restricted materials. Probes and GSM8K generation need Apple silicon with MLX and about 89 GB of disk; E1's tokenizers are rebuilt by E1's fetch script.
- Suggested protocol
- Scripts 01 to 09 in experiments/E02-digit-policy, in order.
Selected exact hypotheses and premises
Premise · P20
85dfefb0-eb84-4deb-b93c-7b494a10b283Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.
Reason for this revision
Initial proposal.