Mediated share
The average causal mediation effect of a mediator divided by the total effect of a treatment on an outcome, with an item bootstrap interval.
Hypothesis
Sign in with GitHubHypothesis · H50 · Author-curated prediction
Bielik-PL-11B-v3.0-Instruct against Bielik-11B-v3.0-Instruct, and the two arms of Qwen2.5-1.5B after 500,170,752 Qwen tokens of continued pretraining on the same document sequence; words matched on frequency decile and character length, with within-type variation across contexts; identity specificity required: patching a different word of the same token length must change the emission to that word. Polish texts are the comparison on which the mediated share is predicted near zero, without a fixed bound.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 1.7686 on English Python code (95% interval 1.7260 to 1.8066), above its 1.5466 on the English preamble.APT4's fertility tax on English Python code.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on Python code is 0.5941 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 0.3795 for Qwen2.5-1.5B with its own tokenizer.The Python residual of the APT4 arm at matched continued pretraining.
finding · premise
After 0.5B tokens of the same continued pretraining, bits per byte on math_clean arithmetic statements is 1.7672 for Qwen2.5-1.5B with APT4 by Fast Vocabulary Transfer and 1.5157 for Qwen2.5-1.5B with its own tokenizer.The math_clean residual of the APT4 arm at matched continued pretraining.
Loading research…
The average causal mediation effect of a mediator divided by the total effect of a treatment on an outcome, with an item bootstrap interval.
In the pair of pair_1_subject and pair_1_comparator and in the pair of pair_2_subject and pair_2_comparator, the measure of the mediator in the effect of the treatment on the outcome is at least threshold on text_1 and on text_2.
The shallowest layer at which a model's hidden state at the final token of a word in context, patched into a fixed neutral prompt that asks for a word to be repeated, makes the model emit that word's surface form.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Qwen2.5-1.5B after continued pretraining with its original tokenizer on a fixed document sequence of about 80% Polish FineWeb2-HQ and 20% English SlimPajama-6B by Qwen tokens, without a mathematics or code slice.
Qwen2.5-1.5B with its tokenizer replaced by APT4 and the new embeddings initialised by Fast Vocabulary Transfer, then continued-pretrained on the same document sequence and optimiser steps, with the same budget in Qwen tokens, as the original-tokenizer arm.
331 Python files from GitHub, truncated at 700 words.
Synthetic English arithmetic, comparison, sorting and unit-conversion statements with their answers, derived from generated arithmetic probes.
{
"wording": "On English Python code and math_clean statements, word reconstruction depth mediates at least 30% of the per-token surprisal penalty of the APT4 model over its original-tokenizer counterpart, in the Bielik 11B v3 pair and in the Qwen2.5-1.5B continued-pretraining pair.",
"predicate": {
"type": "concept",
"key": "mediates_at_least"
},
"roles": [
{
"role": "measure",
"definition": "Quantity bounded.",
"value": {
"type": "concept",
"key": "mediated_share"
}
},
{
"role": "mediator",
"definition": "Variable carrying the effect.",
"value": {
"type": "concept_ref",
"versionId": "02ed29ad-dcef-43db-a25a-c9914e513112",
"key": "reconstruction_depth"
}
},
{
"role": "treatment",
"definition": "Variable whose effect is decomposed.",
"value": {
"type": "text",
"value": "difference in tokens per word between the pair's tokenizers"
}
},
{
"role": "outcome",
"definition": "Variable affected.",
"value": {
"type": "text",
"value": "per-token surprisal penalty of the subject over the comparator"
}
},
{
"role": "threshold",
"definition": "Lower bound on the measure.",
"value": {
"type": "decimal",
"value": "0.3",
"unit": "ratio"
}
},
{
"role": "pair_1_subject",
"definition": "APT4 model of the first pair.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "pair_1_comparator",
"definition": "Original-tokenizer model of the first pair.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "pair_2_subject",
"definition": "APT4 model of the second pair.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "apt4_fvt_cpt_arm"
}
},
{
"role": "pair_2_comparator",
"definition": "Original-tokenizer model of the second pair, after the same continued pretraining.",
"value": {
"type": "concept_ref",
"versionId": "9a53082a-65e8-4c6a-82be-33dda5810584",
"key": "original_tokenizer_cpt_arm"
}
},
{
"role": "text_1",
"definition": "First text.",
"value": {
"type": "concept_ref",
"versionId": "4acc878a-a4a3-45bd-bb03-2cd8541a7856",
"key": "english_python_code"
}
},
{
"role": "text_2",
"definition": "Second text.",
"value": {
"type": "concept_ref",
"versionId": "1c1e659f-9adf-49a7-8976-649eb293aaa7",
"key": "math_clean_statements"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Fertility and vocabulary allocation. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →