Finding

Sign in with GitHub
← Publications

Finding · P250 · Author-curated

None of the 12 model, language and benchmark cells has a 95% Wilson interval for likelihood multiple-choice accuracy that includes the chance level 0.25; the lowest lower bound is 0.3057.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Cells including chance

metric
Likelihood multiple-choice accuracyQuantity measured per cell.
count including
0 cellsCells whose interval includes the chance level.
count total
12 cellsCells checked.
chance level
0.25 accuracyAccuracy of uniform guessing among four options.
lowest interval low
0.3057 accuracySmallest lower bound of the cells' 95% Wilson intervals.
cells
Qwen2.5-1.5B and its two 500M-token continued-pretraining checkpoints; English and Polish; translation-paired Belebele and MMLU itemsCells checked.

Experimental provenance

Method and evaluation protocol
Wilson 95% interval of the length-normalised accuracy of each model, language and benchmark cell over the analysed pairs, compared with 0.25.
Dataset
Translation-paired Belebele and MMLU items after the exclusions.Version: unspecified · Access: public
Reported results
0 of 12 cells include 0.25; lowest lower bound 0.3057.
Uncertainty and replication
Wilson score intervals; no multiplicity adjustment.
Limitations
The floor rule is checked on accuracy only; concentrated letter predictions can keep a cell above chance without item knowledge. Letter priors: on Polish MMLU items the checkpoint that kept the original tokenizer predicts A for 4 of 600 items and the APT4 checkpoint predicts B for 437, against 134 to 160 gold items per letter; on Polish Belebele items the APT4 checkpoint predicts B for 360 and the other checkpoint D for 326 of 600; the base model's predictions are the least concentrated (all scored rows, results/metrics.json pred_count_*).

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Cells including chance

count_including of the count_total evaluation cells have a 95% Wilson interval for the metric that includes chance_level; lowest_interval_low is the smallest lower bound among the cells.

Key cells_including_chance · version 70d37332-4846-4767-bb25-ae8a5ab6d2c2

Likelihood multiple-choice accuracy

Fraction of items for which the option letter with the highest log-probability per continuation token, scored zero-shot after an answer scaffold in the item's language, is the gold letter.

Key paired_mcq_accuracy · version 50041e38-5826-4b54-b39b-fdff21892c61