At most on each
For model_1 and for model_2, the measure of the metric over continuations from fork_point of continuation_tokens tokens is at most threshold on each of the texts.
Hypothesis
Sign in with GitHubHypothesis · H53 · Author-curated prediction
Three continuations per arm; the spread is a lower bound on the variance of full training runs.
finding · premise
On Python code, the bits per byte of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer exceed those of Qwen2.5-1.5B with its own tokenizer by 0.5130, 0.2758, 0.2466, 0.2341 and 0.2145 after 100M, 200M, 300M, 400M and 500M tokens of the same continued pretraining.A formal-text gap reported from one training run per arm.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on math_clean arithmetic statements is 1.7672 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 1.5157 for Qwen2.5-1.5B with its own tokenizer.A formal-text gap reported from one training run per arm.
Loading research…
For model_1 and for model_2, the measure of the metric over continuations from fork_point of continuation_tokens tokens is at most threshold on each of the texts.
Standard deviation of a metric over continuations of one training state that differ only in the data order of their tokens.
Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
{
"wording": "The standard deviation of bits per byte over three continuations of each continued-pretraining arm of Qwen2.5-1.5B from a shared state at 450M tokens, differing only in the data order of their final 50M tokens, is at most 0.01 on each of nine evaluation texts other than GSM8K problems.",
"predicate": {
"type": "concept",
"key": "at_most_on_each"
},
"roles": [
{
"role": "measure",
"definition": "Quantity bounded.",
"value": {
"type": "concept",
"key": "between_continuation_sd"
}
},
{
"role": "metric",
"definition": "Metric whose variation is measured.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "bits_per_byte"
}
},
{
"role": "model_1",
"definition": "First arm.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "model_2",
"definition": "Second arm.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "fork_point",
"definition": "Continued-pretraining tokens of the shared state.",
"value": {
"type": "decimal",
"value": "450000000",
"unit": "tokens"
}
},
{
"role": "continuation_tokens",
"definition": "Tokens of each continuation.",
"value": {
"type": "decimal",
"value": "50000000",
"unit": "tokens"
}
},
{
"role": "texts",
"definition": "Texts scored.",
"value": {
"type": "text",
"value": "the Polish FineWeb2-HQ holdout, Polish Wikipedia science articles, Polish PES examination questions, Polish reviews, the English SlimPajama holdout, English arXiv abstracts, English LaTeX method sections, English Python code and math_clean statements"
}
},
{
"role": "threshold",
"definition": "Upper bound on the measure.",
"value": {
"type": "decimal",
"value": "0.01",
"unit": "bits per byte"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →