← HypothesesTwo 11B instruction-tuned models under one frozen tool-loop protocol on synthetic tasks; a gap measures the tokenizer change, continued pretraining and post-training together.
Exact premises and relationships
Structured prediction and concept definitions
100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.
The observed value of the metric for the subject and the comparator exceeds the value the prediction gives for the same task cells.
The comparator's end-to-end success rate minus the subject's over the same task cells. A task cell is one task instance with one instruction language; it succeeds when the trajectory ends with a correct FINAL value, or a solution that passes the hidden tests, within the step budget.
Predicted end-to-end success of a model on a template in one language: the product over the template's annotated steps of 1 - (1 - p)^(1 + r), where p is the model's one-shot success rate on the probe for the step's type in that language and r = floor((8 - number of steps) / number of steps). The predicted gap is the comparator's mean predicted success minus the subject's.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
{
"wording": "On multi-step tool-use tasks, the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct exceeds the gap predicted by composing the two models' one-shot per-step success rates.",
"predicate": {
"type": "concept",
"key": "exceeds_prediction"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "end_to_end_success_gap"
}
},
{
"role": "subject",
"definition": "Model after the tokenizer change.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "comparator",
"definition": "Model before the tokenizer change.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "benchmark",
"definition": "Tasks the metric is measured on.",
"value": {
"type": "concept",
"key": "agentic_task_suite"
}
},
{
"role": "prediction",
"definition": "Prediction the observed value is compared with.",
"value": {
"type": "concept",
"key": "step_composition_prediction"
}
},
{
"role": "setting",
"definition": "Conditions of the measurement.",
"value": {
"type": "text",
"value": "ReAct tool loop with a persistent Python sandbox, at most 8 steps per task, greedy decoding"
}
}
]
}Correct, supersede, retract or dispute
Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Agentic compounding. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →