Finding

Sign in with GitHub
← Publications

Finding · P37 · Author-curated

On the 102 games of E2 (ZendoBench 1.0.0 dev, the first 2 items of each family), Qwen3.5-4B (MLX) forced to run 16 experiments before it may submit contradicts evidence it has already been shown in 97 of its 189 submissions (0.513), above H2's pre-registered threshold of 0.10; without an intervention (E1's A8 on the same games) the share is 0.217 (44 of 203). E2 attempt A10.

Published by @stw2 · 2026-10-08 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Evaluation estimate

subject
Qwen3.5-4BThe model evaluated.
setting
E2 setting qwen3.5-4b-mlx-force16The setting measured.
intervention
Forced experimenting (16)What E2 changed against E1's setting.
evaluation items
ZendoBench 1.0.0 dev, 2 per familyThe games played.
metric
Contradicting submission shareThe quantity estimated.
submissions
189 submissionsSubmissions over which the share is taken (all 102 games).
value
0.5132 share of submissionsPoint estimate.
threshold
0.1 share of submissionsH2's pre-registered threshold.
baseline setting
E1 setting qwen3.5-4b-mlxThe same model without an intervention on the same games.
baseline value
0.2167 share of submissionsThe share under the baseline setting.
preregistered
trueWhether E2's DESIGN.md named this measure before it was measured.

Experimental provenance

Method and evaluation protocol
E2: `zendo_bench diagnose` (ZendoBench 1.0.0) over the attempt's run file; a submission contradicts when its rule mislabels at least one seed, experiment or counterexample shown before it; scripts/03_measures.py.
Dataset
ZendoBench 1.0.0 dev manifest (Room material M1), its first 2 items of each family (--per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14Version: zendo-bench v1.0.0 (git 46c192e0ea10a5140a33c1280edb97b0127cc68c); manifest sha256 e99344a3c3c9e624bf1c0f5782b266b54e97187cf412d9ff2dbe09d0c1e0a5c2 · Access: public
Reported results
force16: contradicting submissions 97/189 (0.5132); among its wrong submissions 97/163 (0.595). A8 on the same games: 44/203 (0.2167). The threshold's reference (DESIGN.md): every E1 system that ran 10 or more experiments a game stayed below 0.10 over E1's 460 games, from E1's diagnose summaries: GPT-6 Luna 55/630 (0.087), DeepSeek-V4-Flash 14/817 (0.017), Qwen3.8-27B 9/727 (0.012); among their wrong submissions, GPT-6 Luna 55/379 (0.145), DeepSeek-V4-Flash 14/661 (0.021) and Qwen3.8-27B 9/480 (0.019).
Uncertainty and replication
Descriptive shares; no interval. One run.
Evidence references
qwen3.5-4b-mlx-force16.jsonl · restrictedexperiments/E02-small-model-interventions/results/runs/qwen3.5-4b-mlx-force16.jsonl (held by the Room owner; not public; ask the reporter)qwen3.5-4b-mlx-force16.diagnose.md · publichttps://github.com/stw2/zendo-lab/blob/9284883f59dfc88aac67d11519464f28c13570d0/experiments/E02-small-model-interventions/results/scores/qwen3.5-4b-mlx-force16.diagnose.mdmeasures.json · publichttps://github.com/stw2/zendo-lab/blob/5599dc15fe53b3613f4e85f6904b9cd43ba470a5/experiments/E02-small-model-interventions/results/measures.jsongpt-6-luna-openrouter-high.diagnose.md · publichttps://github.com/stw2/zendo-lab/blob/3a25d59f04e9084f4cecf84f124d67c647c5e903/experiments/E01-dev-baseline/results/scores/gpt-6-luna-openrouter-high.diagnose.mddeepseek-v4-flash-together.diagnose.md · publichttps://github.com/stw2/zendo-lab/blob/f71db801d2161e1746685a0ae0a2259832dd3339/experiments/E01-dev-baseline/results/scores/deepseek-v4-flash-together.diagnose.mdqwen3.8-27b-fp8-vllm-h100.diagnose.md · publichttps://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.diagnose.md
Limitations
One run per arm on 102 dev games (92 headline); one model; dev only, no sealed evaluation. More evidence gives a rule more scenes to contradict; the threshold was set from E1 systems that saw a similar amount of evidence, and among wrong submissions alone the forced 4B's share (0.595) is still well above theirs. A8's share on these games (0.217) is already above the threshold, with far less evidence.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Evaluation estimate

Predicate. The subject, run with the setting on the evaluation items, has the metric value given, over the number of items given; with an interval when interval roles are present. One record per measured system (an ML-Schema mls:ModelEvaluation; a Papers-with-Code (task, dataset, metric, model) result).

Key evaluation_estimate · version 7f391bde-3e66-43b6-9a50-d299f7b7e5b2

E1 setting qwen3.5-4b-mlx

ZendoBench 1.0.0 MLX batch driver on Apple M4 Max 128 GB (mlx 0.32.2), bf16, 36 games in flight, sampler seed 0; temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence penalty 1.5; thinking budget 63,487 tokens then a 2,048-token answer allowance (a call's cap 65,536), answers constrained to valid actions; reasoning kept; player model; `bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-4b-mlx`.

Key arm_qwen3_5_4b_mlx · version a1427e75-d571-45f4-9513-1276d692d84b

ZendoBench 1.0.0 dev, 2 per family

Evaluation items. The first 2 items of each family of ZendoBench 1.0.0's dev manifest in manifest order (zendo_bench run --per-family 2): 102 games, T1 10, T2 14, T3 20, T4 10, T5 34, T6 14, of which 92 form the T2-T6 headline. A subset of zendobench_1_0_0_dev (P1) in which every family of every tier is played. E2's games.

Key zendobench_1_0_0_dev_per_family_2 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Forced experimenting (16)

Intervention. Until the player has run 16 experiments, its observation's available_actions is ["experiment"] and its answer grammar is ZendoBench's own experiment-only variant; a sentence in the system prompt states the rule. The gate also opens if one more experiment would leave the remaining submissions unaffordable (budget 30, experiment 1, submission 3, at most 2 submissions).

Key forced_experimenting_16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Contradicting submission share

Metric. The share of a player's submissions whose rule mislabels at least one evidence scene shown before it (seeds, experiments, counterexamples); zendo_bench diagnose field contradicting (n of m submissions), over all scored games including T1.

Key contradicting_submission_share · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

E2 setting qwen3.5-4b-mlx-force16

E1's setting arm_qwen3_5_4b_mlx (P7: ZendoBench 1.0.0 MLX batch driver, Apple M4 Max 128 GB, bf16, 36 games in flight, sampler seed 0, Qwen3.5 thinking sampling, thinking budget 63,487 + 2,048-token answer allowance, answers constrained to valid actions) with forced_experimenting_16, applied by E2's wrapper e02-intervention-v1 (experiments/E02-small-model-interventions/scripts/intervene.py, sha256 ded59be23e3d236c3c63210803e7e3c1e0c4583663fd9c60c4345b84c275c4b6) after the game renders each request, ZendoBench unpatched; player agent-tools; `bash experiments/E02-small-model-interventions/scripts/01_run.sh force16`.

Key arm_qwen3_5_4b_mlx_force16 · version b4fb1d38-4c66-4640-bfd9-d3f864fa4142

Qwen3.5-4B

Qwen3.5-4B open weights (bf16), Hugging Face Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, ZendoBench 1.0.0's pinned 4B checkpoint (shard SHA-256 checked on load).

Key qwen3_5_4b · version a1427e75-d571-45f4-9513-1276d692d84b

Exact references

derived from

4a45cc3f-4aa0-4f61-8c88-164a0a483b54

The forced arm's profile.

supports

5ff5056d-13ab-4c2a-ac45-a54f32d4153b

H2's test: at least 10% of the forced arm's submissions contradict the evidence, and the precondition holds.

derived from

be062a4b-4e6a-4de7-be80-5e98a9ff5053

A8's profile on the same games.