Reached later
measure is larger for each of later_texts than for each of earlier_texts, in the same training run of subject under procedure.
Hypothesis
Sign in with GitHubHypothesis · H16 · Author-curated prediction
APT4 FVT transplant of Qwen2.5-0.5B; checkpoints every 100M tokens up to 1B; stated as a descriptive prediction without a decision rule.
finding · premise
On math_clean arithmetic statements, the bits per byte of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer exceed those of Qwen2.5-1.5B with its own tokenizer by 0.3843, 0.3264, 0.2837, 0.2506 and 0.2515 after 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.During continued pretraining the transplant's math_clean gap to the original-tokenizer model narrows slowly.
finding · premise
On Python code, the bits per byte of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer exceed those of Qwen2.5-1.5B with its own tokenizer by 0.5130, 0.2758, 0.2466, 0.2341 and 0.2145 after 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.The same for Python code.
finding · premise
On the Polish FineWeb2-HQ holdout, the bits per byte of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer exceed those of Qwen2.5-1.5B with its own tokenizer by 0.0957, 0.0465, 0.0305, 0.0201 and 0.0147 after 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.The same for the Polish holdout, where the gap narrows faster.
Loading research…
measure is larger for each of later_texts than for each of earlier_texts, in the same training run of subject under procedure.
Training tokens at the first checkpoint where bits per byte on a text has moved at least 90% of the way from its value at the start of recovery pretraining to its value at the last checkpoint.
Further next-token-prediction training of a pretrained language model on additional text.
A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.
{
"wording": "During recovery pretraining of an APT4 FVT transplant, bits per byte on formal-domain text reaches 90% of its final recovery after more training tokens than bits per byte on Polish and English text.",
"predicate": {
"type": "concept",
"key": "reached_later"
},
"roles": [
{
"role": "subject",
"definition": "Model family recovered.",
"value": {
"type": "concept_ref",
"versionId": "f89e740e-7204-4148-8934-85fbc73f323c",
"key": "apt4_fvt_transplant"
}
},
{
"role": "procedure",
"definition": "Training applied to the subject.",
"value": {
"type": "concept_ref",
"versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
"key": "continued_pretraining"
}
},
{
"role": "measure",
"definition": "Quantity compared between texts.",
"value": {
"type": "concept",
"key": "tokens_to_90pct_recovery"
}
},
{
"role": "later_texts",
"definition": "Texts predicted to reach the threshold later.",
"value": {
"type": "text",
"value": "formal-domain text: math_clean and Python code"
}
},
{
"role": "earlier_texts",
"definition": "Texts predicted to reach the threshold earlier.",
"value": {
"type": "text",
"value": "Polish web text, Polish reviews and English web text"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →