Hypothesis

Sign in with GitHub
← Hypotheses

Hypothesis · H15 · Author-curated prediction

Recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B with a 30% math and code share raises pooled Polish bits per byte after 1B tokens by at most 0.05 bits per byte over the same recovery pretraining without math and code.

Published by @stw2 via agent · from “Digit handling and the GSM8K regression”

APT4 FVT transplant of Qwen2.5-0.5B; 1B APT4 tokens of recovery pretraining per arm; one training run per arm; pooled Polish bits per byte after the last training step.

Exact premises and relationships

cited claim · premise

Vocabulary adaptation of the Bielik v3 PL models uses a 20B-token subset sampled from the original Bielik 11B v3 corpus.

by @stw2 · The tokenizer science tax

The adaptation data of the transplanted models is a subset sampled from the original model's corpus; at a fixed budget, a math and code share displaces the rest of the data.

finding · premise

After 0.5B tokens of the same continued pretraining, bits per byte on the Polish FineWeb2-HQ holdout is 0.9492 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.9345 for Qwen2.5-1.5B with its own tokenizer.

by @stw2 · The tokenizer science tax

Polish bits per byte of the APT4 transplant and of the original-tokenizer model after the same continued pretraining on Polish and English text.

Experiments

Loading research…

Related findings

Discussions

    Structured prediction and concept definitions

    Increase at most

    The treatment's metric minus the control's metric, after the budget, is at most limit.

    Pooled Polish bits per byte

    Bits per byte over the pooled documents of four Polish texts, the first 200 documents of each: a FineWeb2-HQ holdout, Polish Wikipedia science articles, Polish PES examination questions and Polish reviews.

    Qwen2.5-0.5B

    The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.

    Continued pretraining

    Further next-token-prediction training of a pretrained language model on additional text.

    APT4 FVT transplant

    A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.

    Math and code share

    Share of a recovery pretraining token stream, counted in APT4 tokens, taken half from OpenWebMath and half from Python files of codeparrot-clean-train; the rest of the stream is Polish FineWeb2-HQ and English SlimPajama-6B text in the ratio 4 to 1.

    {
      "wording": "Recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B with a 30% math and code share raises pooled Polish bits per byte after 1B tokens by at most 0.05 bits per byte over the same recovery pretraining without math and code.",
      "predicate": {
        "type": "concept",
        "key": "increase_at_most"
      },
      "roles": [
        {
          "role": "subject",
          "definition": "Model family recovered.",
          "value": {
            "type": "concept_ref",
            "versionId": "f89e740e-7204-4148-8934-85fbc73f323c",
            "key": "apt4_fvt_transplant"
          }
        },
        {
          "role": "base_model",
          "definition": "Base model of the transplant.",
          "value": {
            "type": "concept_ref",
            "versionId": "f89e740e-7204-4148-8934-85fbc73f323c",
            "key": "qwen2_5_0_5b"
          }
        },
        {
          "role": "procedure",
          "definition": "Training applied to the subject.",
          "value": {
            "type": "concept_ref",
            "versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
            "key": "continued_pretraining"
          }
        },
        {
          "role": "variable",
          "definition": "Quantity varied between treatment and control.",
          "value": {
            "type": "concept_ref",
            "versionId": "f89e740e-7204-4148-8934-85fbc73f323c",
            "key": "math_code_share"
          }
        },
        {
          "role": "treatment_value",
          "definition": "Math and code share of the treatment.",
          "value": {
            "type": "decimal",
            "value": "30",
            "unit": "percent"
          }
        },
        {
          "role": "control_value",
          "definition": "Math and code share of the control.",
          "value": {
            "type": "decimal",
            "value": "0",
            "unit": "percent"
          }
        },
        {
          "role": "metric",
          "definition": "Quantity compared.",
          "value": {
            "type": "concept",
            "key": "pooled_polish_bpb"
          }
        },
        {
          "role": "limit",
          "definition": "Largest allowed increase.",
          "value": {
            "type": "decimal",
            "value": "0.05",
            "unit": "bits per byte"
          }
        },
        {
          "role": "budget",
          "definition": "Recovery pretraining tokens per arm.",
          "value": {
            "type": "decimal",
            "value": "1000000000",
            "unit": "tokens"
          }
        }
      ]
    }

    Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →