Reuse the defining version and key when the meaning fits your assertion.
Evaluation estimate
Predicate. The subject, run with the setting on the evaluation items, has the metric value given, over the number of items given; with an interval when interval roles are present. One record per measured system (an ML-Schema mls:ModelEvaluation; a Papers-with-Code (task, dataset, metric, model) result).
Key evaluation_estimate · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2
Concept JSON · Defining publication
E1 setting qwen3.5-4b-mlx
ZendoBench 1.0.0 MLX batch driver on Apple M4 Max 128 GB (mlx 0.32.2), bf16, 36 games in flight, sampler seed 0; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence penalty 1.5; thinking budget 63,487 tokens then a 2,048-token answer allowance (a call's cap 65,536), answers constrained to valid actions; reasoning kept; player model; `bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-4b-mlx`.
Key arm_qwen3_5_4b_mlx · version a1427e75-d571-45f4-9513-1276d692d84b
Concept JSON · Defining publication
ZendoBench 1.0.0 dev, 2 per family
Evaluation items. The first 2 items of each family of ZendoBench 1.0.0's dev manifest in manifest order (zendo_bench run --per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14, of which 92 form the T2-T6 headline. A subset of zendobench_1_0_0_dev (P1) in which every family of every tier is played. E2's games.
Key zendobench_1_0_0_dev_per_family_2 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142
Concept JSON · Defining publication
Forced experimenting (16)
Intervention. Until the player has run 16 experiments, its observation's available_actions is ["experiment"] and its answer grammar is ZendoBench's own experiment-only variant; a sentence in the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions).
Key forced_experimenting_16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142
Concept JSON · Defining publication
Contradicting submission share
Metric. The share of a player's submissions whose rule mislabels at least one evidence scene shown before it (seeds, experiments, counterexamples); zendo_bench diagnose field contradicting (n of m submissions), over all scored games including T1.
Key contradicting_submission_share · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142
Concept JSON · Defining publication
E2 setting qwen3.5-4b-mlx-force16
E1's setting arm_qwen3_5_4b_mlx (P7: ZendoBench 1.0.0 MLX batch driver, Apple M4 Max 128 GB, bf16, 36 games in flight, sampler seed 0, Qwen3.5 thinking sampling, thinking budget 63,487 + 2,048-token answer allowance, answers constrained to valid actions) with forced_experimenting_16, applied by E2's wrapper e02-intervention-v1 (experiments/E02-small-model-interventions/scripts/intervene.py, sha256 ded59be23e3d236c3c63210803e7e3c1e0c4583663fd9c60c4345b84c275c4b6) after the game renders each request, ZendoBench unpatched; player agent-tools; `bash experiments/E02-small-model-interventions/scripts/01_run.sh force16`.
Key arm_qwen3_5_4b_mlx_force16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142
Concept JSON · Defining publication
Qwen3.5-4B
Qwen3.5-4B open weights (bf16), Hugging Face Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, ZendoBench 1.0.0's pinned 4B checkpoint (shard SHA-256 checked on load).
Key qwen3_5_4b · version a1427e75-d571-45f4-9513-1276d692d84b
Concept JSON · Defining publication