Finding

Sign in with GitHub
← Publications

Finding · P130 · Author-curated

Untouched Qwen2.5-1.5B has 0.0524 bits per byte on GSM8K problems, the lowest of 10 texts scored; the next lowest is 0.3417 on Python code.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Lowest among texts

metric
Bits per byteQuantity compared.
subject
Qwen2.5-1.5BModel scored.
evaluation text
GSM8K problemsText with the lowest value.
value
0.0524 bits per byteBits per byte on the evaluation text.
runner up text
English Python codeText with the next lowest value.
runner up value
0.3417 bits per byteBits per byte on the runner-up text.
texts compared
10 textsTexts scored.
scope
first 200 documents of each text; 800 math_clean statementsDocuments scored.

Experimental provenance

Method and evaluation protocol
Next-token negative log-likelihood in bits summed over non-overlapping 2048-token windows of each document, the first token of each window unscored, divided by the documents' UTF-8 bytes; bf16 on Apple silicon, 128 GB unified memory; arm-B tokenizers loaded through a whitespace canary and the frozen APT4 reference.
Dataset
Ten texts: Polish and English holdouts, seven E1 corpora and math_clean.Version: unspecified · Access: restricted
Reported results
sci_gsm8k 0.0524; sci_python 0.3417; sci_latex 0.6436; sci_arxiv 0.6935; en 0.8294; pl_wiki_sci 1.0930; pl 1.2043; pl_pes 1.4071; math_clean 1.5023; pl_informal 1.6948
Uncertainty and replication
One scoring pass, repeated once with identical values; no interval.
Limitations
The design and report read the low value as memorisation of GSM8K in Qwen2.5's pretraining data; no decontamination check was run. One run per arm; values at a constant learning rate before the decay anneal the design owed; 1.5B parameters and 0.5B tokens against the paper's 11B and 20B; Fast Vocabulary Transfer instead of FOCUS; the design was committed with the results and fixed no decision rule; the training text was not kept.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Lowest among texts

The subject's metric on the evaluation text is the lowest of the texts compared; runner_up_text has the next lowest value.

Key lowest_among_texts · version 4ac9d8db-adb9-49c5-b0dd-e12ea68e549b

GSM8K problems

The 7,473 questions of the GSM8K training split.

Key english_gsm8k_problems · version 4ab4fa55-c833-4de4-bec3-4d5b68245a5e

Bits per byte

Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.

Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584

English Python code

331 Python files from GitHub, truncated at 700 words.

Key english_python_code · version 4acc878a-a4a3-45bd-bb03-2cd8541a7856

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5

Exact references

related

36da68d9-97cf-4575-ad42-ed56e0e895e9

GSM8K comparisons of models built on continued pretraining.