Finding

Sign in with GitHub
← Publications

Finding · P10 · Author-curated

On ZendoBench 1.0.0 dev (460 games), Qwen3.8-27B-FP8 (vLLM, H100) runs 13.80 experiments per game at 0.45 bits of expected information each, and first submits with a median of 2 rule classes still consistent with the evidence (mean win probability 0.58; 456 games with a submission). E1 attempt A5.

Published by @stw2 · 2026-10-06 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Behaviour profile

subject
Qwen3.8-27B-FP8The model evaluated.
setting
E1 setting qwen3.8-27b-fp8-vllm-h100Backend, hardware, sampling and harness settings of the run.
evaluation items
ZendoBench 1.0.0 devThe games played.
games
460 gamesScored games described (all tiers, T1 included).
experiments per game
13.80 experiments per gameValue of the concept experiments_per_game (defined in the seed-only finding).
expected information gain
0.446 bits per experimentValue of the concept expected_information_gain.
classes alive at first submission
2 rule classesValue of the concept classes_alive_at_first_submission (a median).
win probability at first submission
0.582 probabilityValue of the concept win_probability_at_first_submission (a mean).
games with a submission
456 gamesGames over which the two first-submission statistics are taken.
preregistered
trueWhether DESIGN.md (E1) named these measures before measuring (as descriptive measures).

Experimental provenance

Method and evaluation protocol
E1: `zendo_bench diagnose` (ZendoBench 1.0.0) over the attempt's run files, from each game's engine record alone.
Dataset
ZendoBench 1.0.0 dev manifest: 460 games (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), Room material M1Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
exp_per_game 13.7957, bits_per_exp 0.4658, eig 0.4457, zero-information share 0.4538, alive_first 2.0, p_first 0.5815, submissions per game 1.5804, contradicting submissions 9/727, counterexample uptake 271/271.
Uncertainty and replication
Descriptive; no interval. Medians and means over the stated games.
Evidence references
qwen3.8-27b-fp8-vllm-h100.jsonl · restrictedexperiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.jsonl (held by the Room owner; not public; ask the reporter)qwen3.8-27b-fp8-vllm-h100.part2.jsonl · restrictedexperiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.part2.jsonl (held by the Room owner; not public; ask the reporter)qwen3.8-27b-fp8-vllm-h100.part3.jsonl · restrictedexperiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.part3.jsonl (held by the Room owner; not public; ask the reporter)qwen3.8-27b-fp8-vllm-h100.diagnose.md · publichttps://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.diagnose.md
Limitations
One run per system on the dev split, not sealed. Each system runs at its publisher's recommended sampling and reasoning effort, so settings differ between systems. Dev rule catalogs are public; ZendoBench makes no contamination-free claim. The run spans three boxes: two infrastructure interruptions stopped it, and 60 unfinished games were replayed with --exclude-finished; every game is scored once. Covers all 460 games including T1, while the headline covers T2-T6.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Behaviour profile

Predicate. Descriptive statistics of how the subject plays the evaluation items under the setting, from zendo_bench diagnose (ZendoBench 1.0.0): how much it experiments, the information its experiments gain, and how much uncertainty remains at its first submission. No interval.

Key behaviour_profile · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

ZendoBench 1.0.0 dev

Evaluation items. The 460 games of ZendoBench 1.0.0's dev manifest (T1 30, T2 60, T3 150, T4 60, T5 100, T6 60), file sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 (Room material M1), from github.com/stw2/zendo-bench tag v1.0.0, commit 46c192e. In a game a hidden rule labels scenes of 1 to N pieces; the player builds scenes to have them labelled (experiments, budget 30) and submits rules (at most 2). Close to exact learning from membership and equivalence queries.

Key zendobench_1_0_0_dev · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

E1 setting qwen3.8-27b-fp8-vllm-h100

vLLM 0.30.0 on one H100 80 GB (Daytona, on-demand), FP8 weights, bf16 KV cache, prefix caching off, max-model-len 73,728; temperature 1.0, top_p 0.95, top_k 20, presence penalty 0, reasoning_effort xhigh, max_tokens 65,536; 27 games in flight; reasoning text kept; player model, ZendoBench 1.0.0; `bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.8-27b-fp8-vllm-h100`.

Key arm_qwen3_8_27b_fp8_vllm_h100 · version 9f72911f-8a38-4906-90c6-2bb0d3637236

Qwen3.8-27B-FP8

Qwen3.8-27B-FP8 open weights, Hugging Face Qwen/Qwen3.8-27B-FP8 at revision 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a; every model file checked against the Hub's SHA-256 before serving.

Key qwen3_8_27b_fp8 · version 9f72911f-8a38-4906-90c6-2bb0d3637236

Exact references

related

9f72911f-8a38-4906-90c6-2bb0d3637236

The headline estimate from the same attempt.