Word reconstruction depth
The shallowest layer at which a model's hidden state at the final token of a word in context, patched into a fixed neutral prompt that asks for a word to be repeated, makes the model emit that word's surface form.
Hypothesis
Sign in with GitHubHypothesis · H49 · Author-curated prediction
Bielik-PL-11B-v3.0-Instruct against Bielik-11B-v3.0-Instruct, and the two arms of Qwen2.5-1.5B after 500,170,752 Qwen tokens of continued pretraining on the same document sequence; words matched on frequency decile and character length, with within-type variation across contexts; identity specificity required: patching a different word of the same token length must change the emission to that word.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 1.7686 on English Python code (95% interval 1.7260 to 1.8066), above its 1.5466 on the English preamble.APT4's fertility tax on English Python code.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on Python code is 0.5941 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.3795 for Qwen2.5-1.5B with its own tokenizer.The Python residual of the APT4 arm at matched continued pretraining.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on math_clean arithmetic statements is 1.7672 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 1.5157 for Qwen2.5-1.5B with its own tokenizer.The math_clean residual of the APT4 arm at matched continued pretraining.
Loading research…
The shallowest layer at which a model's hidden state at the final token of a word in context, patched into a fixed neutral prompt that asks for a word to be repeated, makes the model emit that word's surface form.
Over the word_set on text_1 and text_2, matched on the matching variables, the mean metric is larger in pair_1_subject than in pair_1_comparator and larger in pair_2_subject than in pair_2_comparator.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
331 Python files from GitHub, truncated at 700 words.
Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.
{
"wording": "On English Python code and math_clean statements, words that the APT4 tokenizer splits into more pieces than the comparator's tokenizer, matched on frequency decile and character length, have a larger mean reconstruction depth in Bielik-PL-11B-v3.0-Instruct than in Bielik-11B-v3.0-Instruct, and in Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer than in Qwen2.5-1.5B with its own tokenizer after the same continued pretraining.",
"predicate": {
"type": "concept",
"key": "mean_greater_in_each_pair"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "reconstruction_depth"
}
},
{
"role": "pair_1_subject",
"definition": "APT4 model of the first pair.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "pair_1_comparator",
"definition": "Original-tokenizer model of the first pair.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "pair_2_subject",
"definition": "APT4 model of the second pair.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "pair_2_comparator",
"definition": "Original-tokenizer model of the second pair, after the same continued pretraining.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "word_set",
"definition": "Words compared.",
"value": {
"type": "text",
"value": "words that the subject's tokenizer splits into more pieces than the comparator's tokenizer"
}
},
{
"role": "matching",
"definition": "Variables the words are matched on.",
"value": {
"type": "text",
"value": "frequency decile and character length"
}
},
{
"role": "text_1",
"definition": "First text.",
"value": {
"type": "concept_ref",
"versionId": "4acc878a-a4a3-45bd-bb03-2cd8541a7856",
"key": "english_python_code"
}
},
{
"role": "text_2",
"definition": "Second text.",
"value": {
"type": "concept_ref",
"versionId": "1c1e659f-9adf-49a7-8976-649eb293aaa7",
"key": "math_clean_statements"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →