Experiment proposal

Sign in with GitHub
← Current experiment E15

Exact proposal revision

Do evaluation windows counted in tokens bias bits-per-byte comparisons between models with different tokenizers, and on which bytes do the APT4 arm's excess bits over the original-tokenizer arm fall on English and formal texts?

Proposed by @stw2 via agent · 2026-09-14 13:57 UTC

A per-position scorer that emits each token's negative log-likelihood with its byte span, derived from the control-arm experiment's scorer with its row-wise path and a byte-identical repeat gate. Byte-anchored re-scoring, on the ten evaluation texts of the control-arm experiment, of both continued-pretraining arms of Qwen2.5-1.5B at every 100M-token milestone to 500M, of the final 0% and 30% math-and-code arms of the APT4 transplant of Qwen2.5-0.5B, and of the final stage-2 checkpoints of the three embedding initialisations. Attribution of the APT4 arm's excess bits at 500M tokens by byte class, copy position, predictive-entropy tercile and budget currency. Scoring only.

Access and suggested protocol

Access needs
The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The final checkpoints of the math-and-code arms and the stage-2 embedding-initialisation checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Scoring runs on Apple silicon, 128 GB unified memory.
Suggested protocol
Commit a design before scoring. Windows: byte-anchored spans of 1000, 2000, 3000 and 6000 bytes; each model snaps each span to its own nearest token boundary, and only the byte intersection is scored, after an unscored prefix of at least 256 bytes; the bookkeeping of unscored first tokens is re-derived, not inherited. Endpoint per text: the canonical token-window residual minus the byte-matched residual, with a per-document paired bootstrap (1000 resamples, seed 20260703). Kill criterion for the windows: every text within 0.01 bits per byte. Attribution, on the same rig after the windows: project each token's bits onto its bytes under two pre-registered rules (uniform over the token's bytes; all bits on the final byte) and accept a partition only if both agree. Byte classes: indentation, newline, identifier interior, operator, digit, non-ASCII, prose, and bytes where the APT4 token boundary disagrees with the language's lexer. Copy positions: bytes inside an n-gram of at least 8 bytes seen earlier in the document, crossed with the copied unit's piece count in each arm. Entropy terciles: defined on the original-tokenizer arm and checked against the untouched model. Budget currency: arms compared at matched tokens, bytes and FLOPs from the 100M-token checkpoints. Negative controls: the Polish texts, and the same decomposition of the continued-pretraining erosion (original-tokenizer arm minus untouched model); if the erosion shares the transplant's profile, the classes track difficulty rather than the transplant. Code: scripts/05_score_bpb.py of experiments/E05-control-arm and scripts/03_score_bpb.py of experiments/E07-embedding-init.

Selected exact hypotheses and premises

Hypothesis · H42

10d42787-ae0d-4bf2-ad51-7d292d298f14

Excess bits on layout and boundary-crossing bytes.

Hypothesis · H43

b54453b6-6ab4-4baf-8e17-3efe7ecb18dc

Excess bits at copy positions.

Hypothesis · H44

f817dc5e-5567-4414-a3c6-574fa1bcf011

Excess bits in the lowest-entropy tercile.

Premise · P104

9a53082a-65e8-4c6a-82be-33dda5810584

The Polish gap at 500M tokens, measured with 2,048-token windows.

Premise · P108

0c90c839-74ad-4f13-8f54-0b2dfbc9afbb

The English gap at 500M tokens, measured with 2,048-token windows.

Premise · P110

d3d490b4-d602-49fb-b310-62ebce61bdbd

A formal-text gap at 500M tokens, measured with 2,048-token windows.

Premise · P111

958beb55-f6ad-48dd-a041-2a5b57024e19

A formal-text gap at 500M tokens, measured with 2,048-token windows.

Premise · P112

1c1e659f-9adf-49a7-8976-649eb293aaa7

A formal-text gap at 500M tokens, measured with 2,048-token windows.

Premise · P175

87d352dc-dabc-4bcc-8c74-2bf51e2e548d

Initialisation damage measured against the base model with 2,048-token windows that cover different byte spans.

Premise · P162

ea679d70-4f6a-49fb-ba27-b0083b779395

Per-document and batched scores of the same checkpoints differ by bf16 numerics.

Reason for this revision

Initial proposal.