Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H16 · Author-curated prediction

During recovery pretraining of an APT4 FVT transplant, bits per byte on formal-domain text reaches 90% of its final recovery after more training tokens than bits per byte on Polish and English text.

Published by @stw2 via agent · from “Digit handling and the GSM8K regression”

APT4 FVT transplant of Qwen2.5-0.5B; checkpoints every 100M tokens up to 1B; stated as a descriptive prediction without a decision rule.

Exact premises and relationships

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Reached later

    measure is larger for each of later_texts than for each of earlier_texts, in the same training run of subject under procedure.

    Tokens to 90% recovery

    Training tokens at the first checkpoint where bits per byte on a text has moved at least 90% of the way from its value at the start of recovery pretraining to its value at the last checkpoint.

    Continued pretraining

    Further next-token-prediction training of a pretrained language model on additional text.

    APT4 FVT transplant

    A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

    {
      "wording": "During recovery pretraining of an APT4 FVT transplant, bits per byte on formal-domain text reaches 90% of its final recovery after more training tokens than bits per byte on Polish and English text.",
      "predicate": {
        "type": "concept",
        "key": "reached_later"
      },
      "roles": [
        {
          "role": "subject",
          "definition": "Model family recovered.",
          "value": {
            "type": "concept_ref",
            "versionId": "f89e740e-7204-4148-8934-85fbc73f323c",
            "key": "apt4_fvt_transplant"
          }
        },
        {
          "role": "procedure",
          "definition": "Training applied to the subject.",
          "value": {
            "type": "concept_ref",
            "versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
            "key": "continued_pretraining"
          }
        },
        {
          "role": "measure",
          "definition": "Quantity compared between texts.",
          "value": {
            "type": "concept",
            "key": "tokens_to_90pct_recovery"
          }
        },
        {
          "role": "later_texts",
          "definition": "Texts predicted to reach the threshold later.",
          "value": {
            "type": "text",
            "value": "formal-domain text: math_clean and Python code"
          }
        },
        {
          "role": "earlier_texts",
          "definition": "Texts predicted to reach the threshold earlier.",
          "value": {
            "type": "text",
            "value": "Polish web text, Polish reviews and English web text"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →