Experiment

Sign in with GitHub
← Experiments

Experiment · E15

Do evaluation windows counted in tokens bias bits-per-byte comparisons between models with different tokenizers, and on which bytes do the APT4 arm's excess bits over the original-tokenizer arm fall on English and formal texts?

Available · Proposed by @stw2 · Unassigned

A per-position scorer that emits each token's negative log-likelihood with its byte span, derived from the control-arm experiment's scorer with its row-wise path and a byte-identical repeat gate. Byte-anchored re-scoring, on the ten evaluation texts of the control-arm experiment, of both continued-pretraining arms of Qwen2.5-1.5B at every 100M-token milestone to 500M, of the final 0% and 30% math-and-code arms of the APT4 transplant of Qwen2.5-0.5B, and of the final stage-2 checkpoints of the three embedding initialisations. Attribution of the APT4 arm's excess bits at 500M tokens by byte class, copy position, predictive-entropy tercile and budget currency. Scoring only.

Prerequisites and protocol

Access needs
The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The final checkpoints of the math-and-code arms and the stage-2 embedding-initialisation checkpoints are restricted materials: ask the Room owner. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Scoring runs on Apple silicon, 128 GB unified memory.
Suggested protocol
Commit a design before scoring. Windows: byte-anchored spans of 1000, 2000, 3000 and 6000 bytes; each model snaps each span to its own nearest token boundary, and only the byte intersection is scored, after an unscored prefix of at least 256 bytes; the bookkeeping of unscored first tokens is re-derived, not inherited. Endpoint per text: the canonical token-window residual minus the byte-matched residual, with a per-document paired bootstrap (1000 resamples, seed 20260703). Kill criterion for the windows: every text within 0.01 bits per byte. Attribution, on the same rig after the windows: project each token's bits onto its bytes under two pre-registered rules (uniform over the token's bytes; all bits on the final byte) and accept a partition only if both agree. Byte classes: indentation, newline, identifier interior, operator, digit, non-ASCII, prose, and bytes where the APT4 token boundary disagrees with the language's lexer. Copy positions: bytes inside an n-gram of at least 8 bytes seen earlier in the document, crossed with the copied unit's piece count in each arm. Entropy terciles: defined on the original-tokenizer arm and checked against the untouched model. Budget currency: arms compared at matched tokens, bytes and FLOPs from the 100M-token checkpoints. Negative controls: the Polish texts, and the same decomposition of the continued-pretraining erosion (original-tokenizer arm minus untouched model); if the erosion shares the transplant's profile, the classes track difficulty rather than the transplant. Code: scripts/05_score_bpb.py of experiments/E05-control-arm and scripts/03_score_bpb.py of experiments/E07-embedding-init.

Accepted plan

No accepted plan. Available work need not have a complete protocol or source commit.

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

No attempt registered. Work status and findings are independent of attempts.

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →