Ordering holds on at least a count of domains
The metric of first is at most that of second, which is at most that of third, on at least threshold_count of count_total domains.
Hypothesis
Sign in with GitHubHypothesis · H19 · Author-curated prediction
Bits per byte of APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings, and of the base model, on the first 200 documents of nine evaluation domains, before any training.
cited claim · premise
FOCUS represents each token of the target vocabulary as a sparse linear combination of tokens from the original vocabulary, selected by semantic similarity in an auxiliary embedding space.FOCUS combines overlapping tokens by semantic similarity.
cited claim · premise
Frequency-based Vocabulary Transfer initialises token embeddings by aggregating representations of their constituent subword units, guided by frequency statistics.FVT aggregates constituent subword representations.
cited claim · premise
Random initialisation assigns randomly sampled vectors to new tokens, requiring the model to relearn embeddings from scratch and often resulting in slow convergence.Random initialisation relearns embeddings from scratch.
Loading research…
The metric of first is at most that of second, which is at most that of third, on at least threshold_count of count_total domains.
Fast Overlapping Token Combinations Using Sparsemax: an embedding initialisation for a replaced tokenizer's vocabulary.
Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and a new 32,000-row input embedding matrix, tied to the output head, filled by an embedding initialisation.
An embedding initialisation for a replaced tokenizer's vocabulary.
An embedding initialisation for a replaced tokenizer's vocabulary that samples new vectors at random.
{
"wording": "At time zero, bits per byte of the FOCUS-initialised APT4 transplant of Qwen2.5-1.5B is at most that of the FVT-initialised transplant, which is at most that of the random-initialised transplant, on at least 7 of 9 domains.",
"predicate": {
"type": "concept",
"key": "ordering_on_at_least"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "bits_per_byte"
}
},
{
"role": "model",
"definition": "Models measured.",
"value": {
"type": "concept_ref",
"versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
"key": "qwen_apt4_transplant"
}
},
{
"role": "first",
"definition": "Initialisation predicted to have the lowest value.",
"value": {
"type": "concept_ref",
"versionId": "5a078e71-84ff-4039-9e32-7999a5f679f5",
"key": "focus_init"
}
},
{
"role": "second",
"definition": "Initialisation predicted to have the middle value.",
"value": {
"type": "concept_ref",
"versionId": "01087ca0-917c-46f8-a25e-ebc786018d72",
"key": "fvt_init"
}
},
{
"role": "third",
"definition": "Initialisation predicted to have the highest value.",
"value": {
"type": "concept_ref",
"versionId": "cae5a2ea-f055-4e01-bb4b-86bb0b6c0fd3",
"key": "random_init"
}
},
{
"role": "setting",
"definition": "Training state of the models.",
"value": {
"type": "text",
"value": "time zero: no training after embedding initialisation"
}
},
{
"role": "threshold_count",
"definition": "Minimum number of domains on which the ordering holds.",
"value": {
"type": "decimal",
"value": "7",
"unit": "domains"
}
},
{
"role": "count_total",
"definition": "Domains compared.",
"value": {
"type": "decimal",
"value": "9",
"unit": "domains"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Embedding initialisation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →