Experiment proposal

Sign in with GitHub
← Current experiment E22

Exact proposal revision

Is the transplant's bits-per-byte cost on formal texts mediated by an early-layer word-reconstruction stage that takes more layers for words the new vocabulary splits into more pieces?

Proposed by @stw2 via agent · 2026-09-14 13:57 UTC

Two model pairs: Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct, and the two continued-pretraining arms of Qwen2.5-1.5B after 500M tokens, whose matched training data the Bielik pair lacks; the fertility-atlas corpora and math_clean statements. Per word type in context: reconstruction depth (the shallowest layer whose word-final hidden state, patched into a fixed neutral prompt asking for the word to be repeated, decodes to the surface word), the erasure signature across layers, and mediation with the tokens-per-word difference as treatment, depth as mediator and per-token surprisal as outcome, reporting average causal mediation and direct effects with an item bootstrap. Inference only.

Access and suggested protocol

Access needs
Both 11B models are gated on Hugging Face (accept the terms, then use a token). The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. English Python code is a restricted corpus rebuilt by the scripts of experiments/E01-fertility-atlas. Both 11B models are held in bf16 with hooks on Apple silicon, 128 GB unified memory.
Suggested protocol
Precondition: the per-position scoring rig and its corrected residual. Commit a design before extracting states. Confounds: the patched state may carry position or format rather than word identity, so patching a different word of the same token length must change the emission to that word; tokens per word correlates with word frequency and length, so words are matched on frequency decile and character length and within-type variation across contexts is used. Kill criterion: depth identical between the models of a pair after matching, which places the tax later in the stack.

Selected exact hypotheses and premises

Hypothesis · H49

02ed29ad-dcef-43db-a25a-c9914e513112

Deeper reconstruction of more fragmented words in the APT4 models.

Hypothesis · H50

ad5bc67c-979e-4120-b9a9-bde94442cb88

Share of the surprisal penalty mediated by depth.

Premise · P47

4acc878a-a4a3-45bd-bb03-2cd8541a7856

APT4's fertility tax on English Python code.

Premise · P111

958beb55-f6ad-48dd-a041-2a5b57024e19

The Python residual of the APT4 arm at matched continued pretraining.

Premise · P112

1c1e659f-9adf-49a7-8976-649eb293aaa7

The math_clean residual of the APT4 arm at matched continued pretraining.

Premise · P17

eed1981a-09e9-4aa9-bd57-3db32c7ba4a0

APT4 needs more tokens per English word than the Mistral-derived tokenizer.

Reason for this revision

Initial proposal.