Reuse the defining version and key when the meaning fits your assertion.
Qwen3.5-27B-FP8
Qwen3.5-27B-FP8 open weights, Hugging Face Qwen/Qwen3.5-27B-FP8 at revision 97f5941bf617e31c5e237364a8602ce3f03a551a; every model file checked against the Hub's SHA-256 before serving.
Key qwen3_5_27b_fp8 · version c7096cbf-a84e-4c7d-b719-6fea1e79ca98
Concept JSON
E1 setting qwen3.5-27b-fp8-vllm-h100
vLLM 0.30.0 on one H100 80 GB (Daytona, on-demand), FP8 weights, bf16 KV cache, prefix caching off, max-model-len 73,728; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence penalty 1.5, max_tokens 65,536; 27 games in flight; reasoning text kept; player model, ZendoBench 1.0.0; `bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-27b-fp8-vllm-h100`.
Key arm_qwen3_5_27b_fp8_vllm_h100 · version c7096cbf-a84e-4c7d-b719-6fea1e79ca98
Concept JSON
Evaluation estimate
Predicate. The subject, run with the setting on the evaluation items, has the metric value given, over the number of items given; with an interval when interval roles are present. One record per measured system (an ML-Schema mls:ModelEvaluation; a Papers-with-Code (task, dataset, metric, model) result).
Key evaluation_estimate · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
ZendoBench 1.0.0 dev
Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.
Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
Korn-Graubard 95% interval (MOVER)
Interval method. 95% confidence interval from `zendo_bench score` (ZendoBench 1.0.0): per tier, a Korn-Graubard interval with an effective sample size for games clustered by rule class; the equal-weight headline combines the tier intervals by MOVER (score.json ci_method korn-graubard-mover).
Key korn_graubard_mover_95 · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
Win rate, T2-T6 equal weight
Metric. Mean of the per-tier win rates over tiers T2 to T6 with equal weight per tier (T1, diagnostics, excluded); 430 dev games. A game is won when a submitted rule is equivalent to the hidden rule on 1..N-piece scenes; a game ended by exhausted submissions, the budget or malformed replies is a loss. Computed by `zendo_bench score --verify` (ZendoBench 1.0.0), which replays every game against its engine record. Reported as a proportion.
Key win_rate_t2_t6_equal_weight · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication