Exact proposal revision
Does a math and code share in recovery pretraining of an APT4 transplant remove its regression on arithmetic text and digit arithmetic at a fixed token budget, and what does it cost on Polish text?
APT4 FVT transplants of Qwen2.5-0.5B recovery-pretrained on 1B APT4 tokens with 0%, 10% or 30% math and code, the rest Polish and English web text at 4 to 1; evaluation by bits per byte on ten texts and by a synthetic digit probe.
Access and suggested protocol
- Access needs
- One 4090-class GPU, 24 GB, per arm; the APT4 tokenizer repository is gated; one evaluation holdout, four evaluation corpora and the packed training streams are restricted materials.
- Suggested protocol
- In experiments/E06-math-code-rescue: scripts/00_bootstrap_cuda.sh once per arm on its own GPU, then scripts/06_analyze.py.
Selected exact hypotheses and premises
Hypothesis · H16
bc2c559b-2e70-44d8-87ff-a07a94962628H3, reported without a decision rule: recovery order.
Hypothesis · H17
8fb0835e-9768-4edc-b293-0d8d150d2fe8H3, reported without a decision rule: dose response of the rescue fraction.
Premise · P25
d664a095-2689-445e-ac40-a85c6202a460The composition of the adaptation data is not given.
Premise · P117
d9ab6218-cf24-498a-90f1-ebd986fb3508Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on the English holdout.
Premise · P118
ee34e809-5b64-4f37-a761-829e94494bf7Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on arXiv abstracts.
Premise · P119
bbd085f6-467a-47b4-934f-83c089d74d9aContinued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on LaTeX method sections.
Premise · P120
82fd08a5-cca0-4136-a0c3-a146e9d7d14eContinued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on Python code.
Premise · P121
78eef7eb-2581-4f13-8cfb-3094c07dbe55Continued pretraining without math and code, and without a tokenizer replacement, raises bits per byte on math_clean.
Premise · P132
2566f4fe-c7ef-43f0-bb66-0e7ca02bbe3eThe same continued pretraining lowers digit-probe accuracy.
Premise · P108
0c90c839-74ad-4f13-8f54-0b2dfbc9afbbAfter the same continued pretraining, the APT4 transplant trails the original-tokenizer model on the English holdout.
Premise · P109
11728a31-7776-4852-96d3-cb8a28442854After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on arXiv abstracts.
Premise · P110
d3d490b4-d602-49fb-b310-62ebce61bdbdAfter the same continued pretraining, the APT4 transplant trails the original-tokenizer model on LaTeX method sections.
Premise · P111
958beb55-f6ad-48dd-a041-2a5b57024e19After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on Python code.
Premise · P112
1c1e659f-9adf-49a7-8976-649eb293aaa7After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on math_clean.
Premise · P54
2f04e3ac-4c60-437e-93d3-0e5644c55b44APT4 splits every integer into one token per digit.
Premise · P74
b0f7daf6-d1ff-4a89-a255-3c00464a3b96APT4's digit segmentation does not account for the arithmetic accuracy differences of the 11B pair.
Reason for this revision
Initial proposal.