Context comparison with an undefined ratio
On the task in the language at the threshold, the subject's value of the metric is subject_value, a lower bound when subject_lower_bound is true, and the comparator's is comparator_value, which is 0, so the ratio of subject_value to comparator_value is undefined; interval_low and interval_high bound the 95% bootstrap interval of the ratio over the resamples_defined resamples in which it is defined.
Key context_comparison_undefined_ratio · version ccaa3bc9-94b6-4197-9fd6-993bda035bea
Concept JSON
Bielik-11B-v3.0-Instruct
The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.
Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c
Concept JSON · Defining publication
Effective context in characters
Largest number of characters of documents in a prompt at which a model's unconditional accuracy on a task is at least a threshold, interpolated linearly between measured lengths; unconditional accuracy scores every prompt longer than the model's token limit as incorrect.
Key effective_context_characters · version 54782e3d-eed7-4fa0-9b9d-932294a9001a
Concept JSON · Defining publication
Bielik-PL-11B-v3.0-Instruct
The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.
Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04
Concept JSON · Defining publication
Needle retrieval in scientific documents
Task in which one sentence stating a six-character code or a number is inserted at a set depth into scientific documents packed to a character budget, followed by a question asking for that code or number; English documents are LaTeX method sections and arXiv abstracts, Polish documents are Polish Wikipedia science articles and PES examination questions.
Key needle_retrieval · version 54782e3d-eed7-4fa0-9b9d-932294a9001a
Concept JSON · Defining publication