Qwen2.5-0.5B
The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.
Hypothesis
Sign in with GitHubHypothesis · H14 · Author-curated prediction
APT4 FVT transplant of Qwen2.5-0.5B; 1B APT4 tokens of recovery pretraining per arm; one training run per arm; evaluation after the last training step.
cited claim · premise
Vocabulary adaptation of the Bielik v3 PL models uses a 20B-token subset sampled from the original Bielik 11B v3 corpus.The adaptation data of the transplanted models is a sampled subset whose share of mathematics and code is not given.
cited claim · premise
On GSM8K, Bielik-PL-11B-v3.0-Instruct scores 80.97 and Bielik-11B-v3.0-Instruct 85.60.The GSM8K regression of the transplanted 11B model.
cited claim · premise
Frequency-based Vocabulary Transfer initialises token embeddings by aggregating representations of their constituent subword units, guided by frequency statistics.The embedding initialisation of the transplants.
finding · premise
Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens raises bits per byte on math_clean arithmetic statements from 1.5023 to 1.5157.Continued pretraining without a math and code slice raises math_clean bits per byte even without a tokenizer replacement.
finding · premise
Continued pretraining of Qwen2.5-1.5B with its own tokenizer on 0.5B tokens lowers digit-probe accuracy from 0.5867 to 0.5167.The same continued pretraining lowers digit-probe accuracy without a tokenizer replacement.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on math_clean arithmetic statements is 1.7672 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 1.5157 for Qwen2.5-1.5B with its own tokenizer.After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on math_clean.
finding · premise
After 0.5B tokens of the same continued pretraining, digit-probe accuracy is 0.4756 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.5167 for Qwen2.5-1.5B with its own tokenizer, a paired difference of -0.0411 (95% interval -0.0645 to -0.0177, sign-test p 0.0008).After the same continued pretraining, the APT4 transplant trails the original-tokenizer model on the digit probe.
finding · premise
Before any continued pretraining, digit-probe accuracy is 0.0133 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.5867 for untouched Qwen2.5-1.5B.Digit-probe accuracy of the transplant before continued pretraining.
finding · premise
APT4's tokenization contributes to arithmetic accuracy differences between Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct through its handling of the no-break space, not through digit segmentation.APT4's digit segmentation does not account for the arithmetic accuracy differences of the 11B pair.
Loading research…
The Qwen2.5 base language model with 0.5B parameters, Qwen/Qwen2.5-0.5B.
Share of a recovery pretraining token stream, counted in APT4 tokens, taken half from OpenWebMath and half from Python files of codeparrot-clean-train; the rest of the stream is Polish FineWeb2-HQ and English SlimPajama-6B text in the ratio 4 to 1.
The treatment's rescue fraction against the control is at least threshold on metric_1 over evaluation_text_1 and on metric_2, after budget tokens of procedure applied to subject after its tokenizer replacement.
One minus the ratio of the treatment's gap to the control's gap, where a gap is a trained model's value on a metric minus the base model's value, signed so that positive means worse.
A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.
Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.
Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.
Share of correct greedy completions over 300 synthetic arithmetic items (addition, subtraction, comparison and sorting), each in three prompt formats, with 4-shot plain-text prompts, graded by a dual-locale numeric grader.
Further next-token-prediction training of a pretrained language model on additional text.
{
"wording": "Recovery pretraining of an APT4 FVT transplant of Qwen2.5-0.5B with a 30% math and code share closes at least 80% of the gap to the base model left by the same recovery pretraining without math and code, on both math_clean bits per byte and digit-probe accuracy, at a fixed budget of 1B tokens.",
"predicate": {
"type": "concept",
"key": "rescue_at_least"
},
"roles": [
{
"role": "subject",
"definition": "Model family recovered.",
"value": {
"type": "concept",
"key": "apt4_fvt_transplant"
}
},
{
"role": "base_model",
"definition": "Base model the gaps are measured against.",
"value": {
"type": "concept",
"key": "qwen2_5_0_5b"
}
},
{
"role": "procedure",
"definition": "Training applied to the subject.",
"value": {
"type": "concept_ref",
"versionId": "f0470dd9-4d77-4e01-a0a7-f3c1296a5ac2",
"key": "continued_pretraining"
}
},
{
"role": "variable",
"definition": "Quantity varied between treatment and control.",
"value": {
"type": "concept",
"key": "math_code_share"
}
},
{
"role": "treatment_value",
"definition": "Math and code share of the treatment.",
"value": {
"type": "decimal",
"value": "30",
"unit": "percent"
}
},
{
"role": "control_value",
"definition": "Math and code share of the control.",
"value": {
"type": "decimal",
"value": "0",
"unit": "percent"
}
},
{
"role": "measure",
"definition": "Quantity the threshold applies to.",
"value": {
"type": "concept",
"key": "rescue_fraction"
}
},
{
"role": "threshold",
"definition": "Minimum rescue fraction.",
"value": {
"type": "decimal",
"value": "0.8",
"unit": "ratio"
}
},
{
"role": "metric_1",
"definition": "First metric the rescue fraction is computed on.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "bits_per_byte"
}
},
{
"role": "evaluation_text_1",
"definition": "Text of the first metric.",
"value": {
"type": "concept_ref",
"versionId": "1c1e659f-9adf-49a7-8976-649eb293aaa7",
"key": "math_clean_statements"
}
},
{
"role": "metric_2",
"definition": "Second metric the rescue fraction is computed on.",
"value": {
"type": "concept_ref",
"versionId": "fbf37827-4b31-402b-8eba-0bb9d16d62c9",
"key": "digit_probe_accuracy"
}
},
{
"role": "budget",
"definition": "Recovery pretraining tokens per arm.",
"value": {
"type": "decimal",
"value": "1000000000",
"unit": "tokens"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Digit handling and the GSM8K regression. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →