Ratio exceeds
The numerator model's metric divided by the denominator model's metric, on the task in the language at the threshold, is greater than bound; expected is the value predicted beforehand, where given.
Hypothesis
Sign in with GitHubHypothesis · H11 · Author-curated prediction
Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct, 32,768-token window, greedy decoding; prompts longer than 32,704 tokens are not run and score incorrect.
cited claim · premise
On representative Polish text, the reduction of fertility from 3.22 to 1.62 tokens per word nearly doubles the effective Polish context capacity.The claim of nearly doubled effective Polish context capacity, tested as a capability.
cited claim · premise
On the Polish text of the Constitution preamble, APT4 has a fertility ratio of 1.62 tokens per word and the Mistral-derived tokenizer 3.22.The preamble fertility values the claim rests on.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 0.6466 on Polish Wikipedia science articles (95% interval 0.6346 to 0.6608), above its 0.5020 on the Polish preamble.APT4's token ratio on Polish Wikipedia science articles, one of the two document sources.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 0.6461 on Polish PES examination questions (95% interval 0.6420 to 0.6501), above its 0.5020 on the Polish preamble.APT4's token ratio on PES examination questions, the other document source.
Loading research…
The numerator model's metric divided by the denominator model's metric, on the task in the language at the threshold, is greater than bound; expected is the value predicted beforehand, where given.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
The Polish language.
Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.
{
"wording": "On needle retrieval over Polish scientific documents, the ratio of Bielik-PL-11B-v3.0-Instruct's effective context in characters to Bielik-11B-v3.0-Instruct's, at unconditional accuracy 0.70, exceeds 1, and is expected to be about 1.55.",
"predicate": {
"type": "concept_ref",
"versionId": "54782e3d-eed7-4fa0-9b9d-932294a9001a",
"key": "ratio_exceeds"
},
"roles": [
{
"role": "metric",
"definition": "Quantity compared.",
"value": {
"type": "concept_ref",
"versionId": "54782e3d-eed7-4fa0-9b9d-932294a9001a",
"key": "effective_context_characters"
}
},
{
"role": "task",
"definition": "Task the metric is measured on.",
"value": {
"type": "concept_ref",
"versionId": "54782e3d-eed7-4fa0-9b9d-932294a9001a",
"key": "needle_retrieval"
}
},
{
"role": "language",
"definition": "Language of the documents.",
"value": {
"type": "concept_ref",
"versionId": "86acabda-0237-47be-8e26-a81500c184aa",
"key": "polish"
}
},
{
"role": "threshold",
"definition": "Accuracy threshold of the metric.",
"value": {
"type": "decimal",
"value": "0.70",
"unit": "accuracy"
}
},
{
"role": "numerator",
"definition": "Model whose metric is divided.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "denominator",
"definition": "Model whose metric divides.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "bound",
"definition": "Value the ratio exceeds.",
"value": {
"type": "decimal",
"value": "1",
"unit": "ratio"
}
},
{
"role": "expected",
"definition": "Ratio predicted beforehand.",
"value": {
"type": "text",
"value": "about 1.55"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →