Within-window accuracy
Accuracy of a model on a task over the prompts that fit its token limit.
Hypothesis
Sign in with GitHubHypothesis · H12 · Author-curated prediction
Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, 32,768-token window, greedy decoding; prompts longer than 32,704 tokens are not run and score incorrect. The design writes the contrast as at most 0 and tests it one-sided against 0, Holm-adjusted with the two capacity ratios.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 1.4810 on English LaTeX method sections (95% interval 1.4623 to 1.5010), below its 1.5466 on the English preamble.APT4 needs more tokens than the Mistral-derived tokenizer for the same English text, so the PL model is deeper into its window.
Loading research…
Accuracy of a model on a task over the prompts that fit its token limit.
The subject's metric minus the comparator's metric, over the same prompts of the task in the language at the setting, is below bound.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The English language.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.
{
"wording": "Where both models' prompts fit their token limits, the model deeper into its token window is less accurate: on needle retrieval over English scientific documents at the largest measured length both models fit, Bielik-PL-11B-v3.0-Instruct's accuracy minus Bielik-11B-v3.0-Instruct's accuracy on the same prompts is below 0.",
"predicate": {
"type": "concept",
"key": "paired_difference_below"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "within_window_accuracy"
}
},
{
"role": "subject",
"definition": "Model whose accuracy is reduced by the comparator's.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "comparator",
"definition": "Model whose accuracy is subtracted.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "task",
"definition": "Task the metric is measured on.",
"value": {
"type": "concept_ref",
"versionId": "54782e3d-eed7-4fa0-9b9d-932294a9001a",
"key": "needle_retrieval"
}
},
{
"role": "language",
"definition": "Language of the documents.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "english"
}
},
{
"role": "setting",
"definition": "Length at which the difference is taken.",
"value": {
"type": "text",
"value": "largest measured length at which both models' prompts fit their token limits"
}
},
{
"role": "bound",
"definition": "Value the difference is below.",
"value": {
"type": "decimal",
"value": "0",
"unit": "accuracy"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →