Share at least
At least share_threshold of the measure of the subject over the comparator, summed over the texts, falls on the byte_class, under every projection rule.
Hypothesis
Sign in with GitHubHypothesis · H44 · Author-curated prediction
Checkpoints after 500,170,752 Qwen tokens of continued pretraining, one run per arm; byte-anchored evaluation windows; the partition must agree under both projection rules; the Polish texts and the continued-pretraining erosion of the original-tokenizer arm are negative controls. Terciles are checked against the untouched model.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on English LaTeX method sections is 0.7979 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.6772 for Qwen2.5-1.5B with its own tokenizer.The LaTeX residual whose bytes are partitioned.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on Python code is 0.5941 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.3795 for Qwen2.5-1.5B with its own tokenizer.The Python residual whose bytes are partitioned.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on math_clean arithmetic statements is 1.7672 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 1.5157 for Qwen2.5-1.5B with its own tokenizer.The math_clean residual whose bytes are partitioned.
Loading research…
At least share_threshold of the measure of the subject over the comparator, summed over the texts, falls on the byte_class, under every projection rule.
Bytes of the original-tokenizer arm's tokens whose next-token predictive entropy under that arm lies in the lowest third of that arm's tokens on the text.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
For each byte of a text, the bits a subject model assigns to it minus the bits a comparator model assigns to it, each model's next-token negative log-likelihood in bits being projected from each token onto that token's bytes by a projection rule.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
English LaTeX method sections, English Python code and math_clean statements.
{
"wording": "On English LaTeX method sections, Python code and math_clean statements, at least 60% of the excess bits of Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer over Qwen2.5-1.5B with its own tokenizer, after the same 500M tokens of continued pretraining, fall in the lowest predictive-entropy tercile.",
"predicate": {
"type": "concept",
"key": "share_at_least"
},
"roles": [
{
"role": "measure",
"definition": "Quantity partitioned.",
"value": {
"type": "concept_ref",
"versionId": "10d42787-ae0d-4bf2-ad51-7d292d298f14",
"key": "excess_bits"
}
},
{
"role": "subject",
"definition": "Model whose bits are the minuend.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "comparator",
"definition": "Model whose bits are the subtrahend.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "continued_pretraining_tokens",
"definition": "Tokens of continued pretraining of both arms at the checkpoints named.",
"value": {
"type": "decimal",
"value": "500170752",
"unit": "tokens"
}
},
{
"role": "texts",
"definition": "Texts scored.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "formal_domains"
}
},
{
"role": "byte_class",
"definition": "Bytes predicted to carry the excess.",
"value": {
"type": "concept",
"key": "lowest_entropy_tercile"
}
},
{
"role": "share_threshold",
"definition": "Lower bound on the share of the summed measure.",
"value": {
"type": "decimal",
"value": "0.6",
"unit": "ratio"
}
},
{
"role": "projection_rules",
"definition": "Rules projecting a token's bits onto its bytes.",
"value": {
"type": "text",
"value": "uniform over the token's bytes; all bits on the token's final byte"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →