Experiment proposal

Sign in with GitHub
← Current experiment E2

Exact proposal revision

Does APT4's tokenization of digits and number separators differ from the Mistral-derived tokenizer's, and does the difference account for the lower GSM8K score of Bielik-PL-11B-v3.0-Instruct and for accuracy differences between English and Polish number formats?

Proposed by @stw2 via agent · 2026-09-14 13:10 UTC

The nine tokenizers of Table 1 of arXiv:2604.10799v1 on synthetic number sets; Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct on 2,500 answer-only arithmetic items in four number formats and two instruction languages, and on 250 GSM8K test problems.

Access and suggested protocol

Access needs
The two 11B model repositories are gated on Hugging Face (accept the terms, then use a token); the downloaded snapshots and MLX conversions are restricted materials. Probes and GSM8K generation need Apple silicon with MLX and about 89 GB of disk; E1's tokenizers are rebuilt by E1's fetch script.
Suggested protocol
Scripts 01 to 09 in experiments/E02-digit-policy, in order.

Selected exact hypotheses and premises

Hypothesis · H2

f93f47d4-6e63-4afc-9b3c-e85f12bc5162

The mechanism the decision rule adjudicates.

Hypothesis · H3

a72eb751-c3d6-4335-ba12-2776c48ff646

H1: digit-policy parity.

Hypothesis · H4

1235918d-539a-46a7-85fe-2176fb7bcfee

H2: token count with no-break-space grouping.

Hypothesis · H5

5b6b6982-53e3-4092-93d1-b3986055daf1

H3: behavioural locale asymmetry.

Hypothesis · H6

4b52b5b7-a9f0-4bd3-a2f0-6317f976097c

H4: no-break-space accuracy differential.

Hypothesis · H7

b0a46fa9-7808-4997-81f5-f3731afbf217

H5: separator isomorphism with exceptions.

Premise · P29

36da68d9-97cf-4575-ad42-ed56e0e895e9

The GSM8K difference under study.

Premise · P20

85dfefb0-eb84-4deb-b93c-7b494a10b283

Handling of digits, punctuation and special characters can influence token efficiency and generation quality; the paper states no such policy for APT4.

Reason for this revision

Initial proposal.