Ratio exceeds
The numerator model's metric divided by the denominator model's metric, on the task in the language at the threshold, is greater than bound; expected is the value predicted beforehand, where given.
Hypothesis
Sign in with GitHubHypothesis · H10 · Author-curated prediction
Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, 32,768-token window, greedy decoding; prompts longer than 32,704 tokens are not run and score incorrect.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 1.4810 on English LaTeX method sections (95% interval 1.4623 to 1.5010), below its 1.5466 on the English preamble.APT4's token ratio on English LaTeX method sections, one of the two document sources.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 1.6226 on English arXiv abstracts (95% interval 1.6140 to 1.6308), above its 1.5466 on the English preamble.APT4's token ratio on English arXiv abstracts, the other document source.
Loading research…
The numerator model's metric divided by the denominator model's metric, on the task in the language at the threshold, is greater than bound; expected is the value predicted beforehand, where given.
Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.
Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
The English language.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
{
"wording": "On needle retrieval over English scientific documents, the ratio of Bielik-11B-v3.0-Instruct's effective context in characters to Bielik-PL-11B-v3.0-Instruct's, at unconditional accuracy 0.70, exceeds 1, and is expected to be about 1.5 to 1.6.",
"predicate": {
"type": "concept",
"key": "ratio_exceeds"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept",
"key": "effective_context_characters"
}
},
{
"role": "task",
"definition": "Task the metric is measured on.",
"value": {
"type": "concept",
"key": "needle_retrieval"
}
},
{
"role": "language",
"definition": "Language of the documents.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "english"
}
},
{
"role": "threshold",
"definition": "Accuracy threshold of the metric.",
"value": {
"type": "decimal",
"value": "0.70",
"unit": "accuracy"
}
},
{
"role": "numerator",
"definition": "Model whose metric is divided.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "denominator",
"definition": "Model whose metric divides.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "bound",
"definition": "Value the ratio exceeds.",
"value": {
"type": "decimal",
"value": "1",
"unit": "ratio"
}
},
{
"role": "expected",
"definition": "Ratio predicted beforehand.",
"value": {
"type": "text",
"value": "about 1.5 to 1.6"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →