Finding

Sign in with GitHub
← Publications

Finding · P251 · Author-curated

Two runs of the paired likelihood scoring of Qwen2.5-1.5B on the first 10 items of each of the four item files produce byte-identical score files of 40 rows.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Byte-identical repeat

procedure
likelihood multiple-choice scoring of the first 10 items of each item file, torch MPS bf16Procedure repeated.
subject
Qwen2.5-1.5BModel scored.
runs
2 runsNumber of runs compared.
rows per run
40 rowsRows in each score file.

Experimental provenance

Method and evaluation protocol
Run the scoring script twice with the same arguments and compare the sha256 of the two score files.
Dataset
First 10 items of each of the four translation-paired item files (Belebele and MMLU, English and Polish).Version: unspecified · Access: public
Reported results
Both files sha256 7f3234fd173d904874e4e36e03d27cebdcd2dc74f636b6f93aafae096b79689d.
Uncertainty and replication
Exact byte comparison.
Limitations
The design's smoke gate names all three models; it ran on the base model only. Byte identity on one machine says nothing about other hardware.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Byte-identical repeat

runs repeated runs of the procedure on the subject produce output files with identical bytes (identical sha256), each holding rows_per_run rows.

Key byte_identical_repeat · version 17e3db6d-8386-4e44-a153-2d0d6edde538

Qwen2.5-1.5B

The 1.5B-parameter base language model of the Qwen2.5 series, without further training.

Key qwen2_5_1_5b · version b7e7c5ce-2805-43d7-9f89-ccda038df1b5