Exact proposal revision
Do evaluation windows counted in tokens bias bits-per-byte comparisons between models with different tokenizers, and on which bytes do the APT4 arm's excess bits over the original-tokenizer arm fall on English and formal texts?
A per-position scorer that emits each token's negative log-likelihood with its byte span, derived from the control-arm experiment's scorer with its row-wise path and a byte-identical repeat gate. Byte-anchored re-scoring, on the ten evaluation texts of the control-arm experiment, of both continued-pretraining arms of Qwen2.5-1.5B at every 100M-token milestone to 500M, of the final 0% and 30% math-and-code arms of the APT4 transplant of Qwen2.5-0.5B, and of the final stage-2 checkpoints of the three embedding initialisations. Attribution of the APT4 arm's excess bits at 500M tokens by byte class, copy position, predictive-entropy tercile and budget currency. Scoring only.
Access and suggested protocol
- Access needs
- The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The final checkpoints of the math-and-code arms and the stage-2 embedding-initialisation checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Scoring runs on Apple silicon, 128 GB unified memory.
- Suggested protocol
- Commit a design before scoring. Windows: byte-anchored spans of 1000, 2000, 3000 and 6000 bytes; each model snaps each span to its own nearest token boundary, and only the byte intersection is scored, after an unscored prefix of at least 256 bytes; the bookkeeping of unscored first tokens is re-derived, not inherited. Endpoint per text: the canonical token-window residual minus the byte-matched residual, with a per-document paired bootstrap (1000 resamples, seed 20260703). Kill criterion for the windows: every text within 0.01 bits per byte. Attribution, on the same rig after the windows: project each token's bits onto its bytes under two pre-registered rules (uniform over the token's bytes; all bits on the final byte) and accept a partition only if both agree. Byte classes: indentation, newline, identifier interior, operator, digit, non-ASCII, prose, and bytes where the APT4 token boundary disagrees with the language's lexer. Copy positions: bytes inside an n-gram of at least 8 bytes seen earlier in the document, crossed with the copied unit's piece count in each arm. Entropy terciles: defined on the original-tokenizer arm and checked against the untouched model. Budget currency: arms compared at matched tokens, bytes and FLOPs from the 100M-token checkpoints. Negative controls: the Polish texts, and the same decomposition of the continued-pretraining erosion (original-tokenizer arm minus untouched model); if the erosion shares the transplant's profile, the classes track difficulty rather than the transplant. Code: scripts/05_score_bpb.py of experiments/E05-control-arm and scripts/03_score_bpb.py of experiments/E07-embedding-init.
Selected exact hypotheses and premises
Hypothesis · H42
10d42787-ae0d-4bf2-ad51-7d292d298f14Excess bits on layout and boundary-crossing bytes.
Premise · P104
9a53082a-65e8-4c6a-82be-33dda5810584The Polish gap at 500M tokens, measured with 2,048-token windows.
Premise · P108
0c90c839-74ad-4f13-8f54-0b2dfbc9afbbThe English gap at 500M tokens, measured with 2,048-token windows.
Premise · P110
d3d490b4-d602-49fb-b310-62ebce61bdbdA formal-text gap at 500M tokens, measured with 2,048-token windows.
Premise · P111
958beb55-f6ad-48dd-a041-2a5b57024e19A formal-text gap at 500M tokens, measured with 2,048-token windows.
Premise · P112
1c1e659f-9adf-49a7-8976-649eb293aaa7A formal-text gap at 500M tokens, measured with 2,048-token windows.
Premise · P175
87d352dc-dabc-4bcc-8c74-2bf51e2e548dInitialisation damage measured against the base model with 2,048-token windows that cover different byte spans.
Premise · P162
ea679d70-4f6a-49fb-ba27-b0083b779395Per-document and batched scores of the same checkpoints differ by bf16 numerics.
Reason for this revision
Initial proposal.