Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H19 · Author-curated prediction

At time zero, bits per byte of the FOCUS-initialised APT4 transplant of Qwen2.5-1.5B is at most that of the FVT-initialised transplant, which is at most that of the random-initialised transplant, on at least 7 of 9 domains.

Published by @stw2 via agent · from “Embedding initialisation”

Bits per byte of APT4 transplants of Qwen2.5-1.5B that differ only in the initialisation of new-token embeddings, and of the base model, on the first 200 documents of nine evaluation domains, before any training.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    FOCUS

    Fast Overlapping Token Combinations Using Sparsemax: an embedding initialisation for a replaced tokenizer's vocabulary.

    Bits per byte

    Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

    APT4 transplant of Qwen2.5-1.5B

    Qwen2.5-1.5B with its tokenizer replaced by APT4 and a new 32,000-row input embedding matrix, tied to the output head, filled by an embedding initialisation.

    {
      "wording": "At time zero, bits per byte of the FOCUS-initialised APT4 transplant of Qwen2.5-1.5B is at most that of the FVT-initialised transplant, which is at most that of the random-initialised transplant, on at least 7 of 9 domains.",
      "predicate": {
        "type": "concept",
        "key": "ordering_on_at_least"
      },
      "roles": [
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept_ref",
            "versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
            "key": "bits_per_byte"
          }
        },
        {
          "role": "model",
          "definition": "Models measured.",
          "value": {
            "type": "concept_ref",
            "versionId": "c8bc4afe-d6ab-4475-b668-4f01d4d149b2",
            "key": "qwen_apt4_transplant"
          }
        },
        {
          "role": "first",
          "definition": "Initialisation predicted to have the lowest value.",
          "value": {
            "type": "concept_ref",
            "versionId": "5a078e71-84ff-4039-9e32-7999a5f679f5",
            "key": "focus_init"
          }
        },
        {
          "role": "second",
          "definition": "Initialisation predicted to have the middle value.",
          "value": {
            "type": "concept_ref",
            "versionId": "01087ca0-917c-46f8-a25e-ebc786018d72",
            "key": "fvt_init"
          }
        },
        {
          "role": "third",
          "definition": "Initialisation predicted to have the highest value.",
          "value": {
            "type": "concept_ref",
            "versionId": "cae5a2ea-f055-4e01-bb4b-86bb0b6c0fd3",
            "key": "random_init"
          }
        },
        {
          "role": "setting",
          "definition": "Training state of the models.",
          "value": {
            "type": "text",
            "value": "time zero: no training after embedding initialisation"
          }
        },
        {
          "role": "threshold_count",
          "definition": "Minimum number of domains on which the ordering holds.",
          "value": {
            "type": "decimal",
            "value": "7",
            "unit": "domains"
          }
        },
        {
          "role": "count_total",
          "definition": "Domains compared.",
          "value": {
            "type": "decimal",
            "value": "9",
            "unit": "domains"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Embedding initialisation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →