Values by share on two texts
After budget tokens of recovery pretraining of subject, metric on text_1 is text_1_00, text_1_10 and text_1_30 and on text_2 is text_2_00, text_2_10 and text_2_30 for math and code shares of 0%, 10% and 30%.
Key values_by_share_on_two_texts · version 88d16551-a5e8-4c4d-91e9-503d03b1414e
Concept JSON
Bits per byte
Summed next-token negative log-likelihood of a text in bits divided by the text's UTF-8 byte count.
Key bits_per_byte · version 9a53082a-65e8-4c6a-82be-33dda5810584
Concept JSON · Defining publication
APT4 FVT transplant
A pretrained Qwen2.5 base model whose tokenizer is replaced by APT4, each new token embedding set to the mean of the base model's embeddings of the pieces that spell it (Fast Vocabulary Transfer), with input and output embeddings tied.
Key apt4_fvt_transplant · version f89e740e-7204-4148-8934-85fbc73f323c
Concept JSON · Defining publication
math_clean statements
Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.
Key math_clean_statements · version 1c1e659f-9adf-49a7-8976-649eb293aaa7
Concept JSON · Defining publication
English Python code
331 Python files from GitHub, truncated at 700 words.
Key english_python_code · version 4acc878a-a4a3-45bd-bb03-2cd8541a7856
Concept JSON · Defining publication