Mean prompt-side tokens per trajectory
Mean over trajectories of the prompt and history tokens a model reads, summed over the trajectory's steps.
Hypothesis
Sign in with GitHubHypothesis · H29 · Author-curated prediction
Two 11B instruction-tuned models under one frozen tool-loop protocol on synthetic tasks; a gap measures the tokenizer change, continued pretraining and post-training together.
cited claim · premise
On the official English translation of the Constitution preamble, APT4 has a fertility ratio of 1.98 tokens per word and the Mistral-derived tokenizer 1.28.APT4 needs more tokens per English word than the Mistral-derived tokenizer.
finding · premise
APT4's English fertility tax exceeds its preamble tax on 2 of 4 scientific corpora (English arXiv abstracts and English Python code) and is below it on the other 2 (English LaTeX method sections and GSM8K problems).APT4's English fertility tax on scientific corpora.
finding · premise
Within the 32,704-token prompt limit, needle-retrieval prompts over English scientific documents hold at most 113,897 characters of documents for Bielik-11B-v3.0-Instruct and 77,432 for Bielik-PL-11B-v3.0-Instruct, a ratio of 1.4709.Characters of English documents that fit each model's prompt limit.
finding · premise
On English needle-retrieval and cloze-retrieval prompts, 110 of 322 generations of Bielik-11B-v3.0-Instruct (0.3416) and 0 of 231 generations of Bielik-PL-11B-v3.0-Instruct contain the <think> tag.The original model's <think> tag on English prompts, which adds to its generated tokens.
Loading research…
Mean over trajectories of the prompt and history tokens a model reads, summed over the trajectory's steps.
The subject's mean of ratio_metric divided by the comparator's exceeds ratio_threshold, and the subject's rate_metric exceeds the comparator's, on the stated benchmark, language and setting.
Share of task cells whose trajectory ends because the next prompt would exceed the context window or because the step budget is used up without an accepted FINAL answer.
100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The English language.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
{
"wording": "On English-instructed multi-step tool-use tasks, Bielik-PL-11B-v3.0-Instruct uses more than 1.10 times the mean prompt-side tokens per trajectory of Bielik-11B-v3.0-Instruct and ends a larger share of trajectories by context overflow or step-budget exhaustion.",
"predicate": {
"type": "concept",
"key": "exceeds_ratio_and_rate"
},
"roles": [
{
"role": "ratio_metric",
"definition": "Quantity whose ratio is compared with the threshold.",
"value": {
"type": "concept",
"key": "mean_prompt_tokens"
}
},
{
"role": "ratio_threshold",
"definition": "Ratio the subject's mean divided by the comparator's must exceed.",
"value": {
"type": "decimal",
"value": "1.10",
"unit": "ratio"
}
},
{
"role": "rate_metric",
"definition": "Rate the subject must exceed.",
"value": {
"type": "concept",
"key": "overflow_or_budget_exhaustion_rate"
}
},
{
"role": "subject",
"definition": "Model after the tokenizer change.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "comparator",
"definition": "Model before the tokenizer change.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "benchmark",
"definition": "Tasks the metrics are measured on.",
"value": {
"type": "concept_ref",
"versionId": "0ea2c3bc-bc27-4395-9371-ce71dbf2c962",
"key": "agentic_task_suite"
}
},
{
"role": "language",
"definition": "Instruction language of the task cells.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "english"
}
},
{
"role": "setting",
"definition": "Conditions of the measurement.",
"value": {
"type": "text",
"value": "ReAct tool loop with a persistent Python sandbox, at most 8 steps per task, greedy decoding"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Agentic compounding. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →