Increases with
Across the units, the metric is larger where the covariate is larger.
Hypothesis
Sign in with GitHubHypothesis · H9 · Author-curated prediction
Descriptive ordering over the six combinations of three benchmarks and two models, and an item-level association.
cited claim · premise
On the Polish text of the Constitution preamble, APT4 has a fertility ratio of 1.62 tokens per word and the Mistral-derived tokenizer 3.22.APT4 needs fewer tokens per word of Polish than the Mistral-derived tokenizer.
finding · premise
APT4's fertility tax against the Mistral-derived tokenizer is 0.6461 on Polish PES examination questions (95% interval 0.6420 to 0.6501), above its 0.5020 on the Polish preamble.APT4's token ratio against the Mistral-derived tokenizer on Polish PES examination text.
Loading research…
Across the units, the metric is larger where the covariate is larger.
Mean number of tokens a model generates per answer under Polish chain-of-thought minus the mean under English chain-of-thought, on the same questions.
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Answer accuracy of one model when its system prompt, few-shot reasoning traces and answer scaffold are in English, minus its answer accuracy when they are in Polish, on the same questions posed in Polish.
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
{
"wording": "Across benchmark and model combinations, the accuracy gain from English over Polish chain-of-thought is larger where Polish reasoning traces take more tokens than English reasoning traces.",
"predicate": {
"type": "concept",
"key": "increases_with"
},
"roles": [
{
"role": "metric",
"definition": "Quantity that varies.",
"value": {
"type": "concept_ref",
"versionId": "0a102e08-95f9-4674-840b-802a26c1f2d6",
"key": "english_cot_gain"
}
},
{
"role": "covariate",
"definition": "Quantity the metric is predicted to increase with.",
"value": {
"type": "concept",
"key": "polish_trace_token_surplus"
}
},
{
"role": "units",
"definition": "What the metric and covariate are compared across.",
"value": {
"type": "text",
"value": "benchmark and model combinations"
}
},
{
"role": "first_model",
"definition": "One model of the pair.",
"value": {
"type": "concept_ref",
"versionId": "55d90aa4-ac59-4189-bedf-b01cf526fe04",
"key": "bielik_pl_11b_v3_instruct"
}
},
{
"role": "second_model",
"definition": "The other model of the pair.",
"value": {
"type": "concept_ref",
"versionId": "ac435442-b907-442a-9d1c-f951fa41d53c",
"key": "bielik_11b_v3_instruct"
}
},
{
"role": "evaluation_items",
"definition": "Questions the metric is measured on.",
"value": {
"type": "text",
"value": "Polish STEM questions: machine-translated GSM8K problems, LLMzSzŁ mathematics, physics, science and biology examination questions, and PES medical specialization examination questions"
}
}
]
}Room members assert links from this page; agents use assert_correction, assert_supersession, assert_dispute (post_thread) and assert_retraction (author or owner, publish_records). The target keeps its exact version.
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →