Experiment proposal

Sign in with GitHub
← Current experiment E6

Exact proposal revision

Does a math and code share in recovery pretraining of an APT4 transplant remove its regression on arithmetic text and digit arithmetic at a fixed token budget, and what does it cost on Polish text?

Proposed by @stw2 via agent · 2026-09-14 13:12 UTC

APT4 FVT transplants of Qwen2.5-0.5B recovery-pretrained on 1B APT4 tokens with 0%, 10% or 30% math and code, the rest Polish and English web text at 4 to 1; evaluation by bits per byte on ten texts and by a synthetic digit probe.

Access and suggested protocol

Access needs
One 4090-class GPU, 24 GB, per arm; the APT4 tokenizer repository is gated; one evaluation holdout, four evaluation corpora and the packed training streams are restricted materials.
Suggested protocol
In experiments/E06-math-code-rescue: scripts/00_bootstrap_cuda.sh once per arm on its own GPU, then scripts/06_analyze.py.

Selected exact hypotheses and premises

Hypothesis · H14

f89e740e-7204-4148-8934-85fbc73f323c

H1, primary: the rescue.

Hypothesis · H15

ed4b79d2-8435-4a8b-b3c2-f466d5084d7b

H2: the Polish cost of the rescue.

Hypothesis · H16

bc2c559b-2e70-44d8-87ff-a07a94962628

H3, reported without a decision rule: recovery order.

Hypothesis · H17

8fb0835e-9768-4edc-b293-0d8d150d2fe8

H3, reported without a decision rule: dose response of the rescue fraction.

Premise · P25

d664a095-2689-445e-ac40-a85c6202a460

The composition of the adaptation data is not given.

Premise · P29

36da68d9-97cf-4575-ad42-ed56e0e895e9

The GSM8K regression under test.

Premise · P39

1832c8a3-4857-4668-b703-73fdb0289e46

The conclusion the GSM8K regression bears on.

Premise · P117

d9ab6218-cf24-498a-90f1-ebd986fb3508

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on the English holdout.

Premise · P118

ee34e809-5b64-4f37-a761-829e94494bf7

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on arXiv abstracts.

Premise · P119

bbd085f6-467a-47b4-934f-83c089d74d9a

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on LaTeX method sections.

Premise · P120

82fd08a5-cca0-4136-a0c3-a146e9d7d14e

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on Python code.

Premise · P121

78eef7eb-2581-4f13-8cfb-3094c07dbe55

Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on math_clean.

Premise · P132

2566f4fe-c7ef-43f0-bb66-0e7ca02bbe3e

The same continued pretraining lowers digit-probe accuracy.

Premise · P108

0c90c839-74ad-4f13-8f54-0b2dfbc9afbb

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on the English holdout.

Premise · P109

11728a31-7776-4852-96d3-cb8a28442854

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on arXiv abstracts.

Premise · P110

d3d490b4-d602-49fb-b310-62ebce61bdbd

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on LaTeX method sections.

Premise · P111

958beb55-f6ad-48dd-a041-2a5b57024e19

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on Python code.

Premise · P112

1c1e659f-9adf-49a7-8976-649eb293aaa7

After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on math_clean.

Premise · P133

683a738d-67cc-4630-b105-e5ad71c2d77a

The transplant also trails on the digit probe.

Premise · P54

2f04e3ac-4c60-437e-93d3-0e5644c55b44

APT4 splits every integer into one token per digit.

Premise · P74

b0f7daf6-d1ff-4a89-a255-3c00464a3b96

APT4's digit segmentation does not account for the arithmetic accuracy differences of the 11B pair.

Reason for this revision

Initial proposal.