Execution attempt

Sign in with GitHub
← Experiment E1 · Where do small and strong models stand on ZendoBench 1.0.0 dev, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?

Execution attempt · A6 · planned

E1 arm qwen3.5-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/zendo-lab @ fac2c1b7381bf8c7a821dff4c32c55fbb2b290b0

Reference checked 2026-10-04 23:27 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-27b-fp8-vllm-h100
Working directory
.
Configuration paths
experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json experiments/E01-dev-baseline/scripts/daytona/serve.sh
Parameters
model=Qwen/Qwen3.5-27B-FP8@97f5941b; backend=vLLM 0.30.0, --batch auto; hardware=1x H100 80GB (Daytona, on-demand); sampling=temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; max_tokens=65536
Environment
Harness on macOS (Python 3.13, ZendoBench 1.0.0) over HTTPS to vLLM 0.30.0 on one on-demand H100 80GB (Daytona sandbox, CUDA 13.0 image)
Output directory
experiments/E01-dev-baseline/results/runs

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

T1.wins
14
T2.wins
15
T3.wins
30
T4.wins
0
T5.wins
7
T6.wins
0
T1.scored
30
T2.scored
60
T3.scored
150
T4.scored
60
T5.scored
100
T6.scored
60
T1.win_rate
0.4666666666666667
T2.win_rate
0.25
T3.win_rate
0.2
T4.win_rate
0
T5.win_rate
0.07
T6.win_rate
0
alive_first
27
diagnose.eig
0.6970935985259261
diagnose.zero
0.12738336713995943
headline_items
430
headline_t2_t6
0.10400000000000001
malformed_rate
0.0767054385413392

and 14 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    E1 arm qwen3.5-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Launched 2026-10-04T23:28:05Z at fac2c1b (clean tree): bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-27b-fp8-vllm-h100, against vLLM 0.30.0 on one on-demand H100 (model files matched the Hub's hashes; chat check passed); 27 games in flight (--batch auto). Guard: stops at $200 or 34 h.

    reported · received · @stw2 via agent · attempt only

  3. Progress report

    #3

    01:14Z: 102/460 games finished (T1 30, T2 60, T3 12), 669 calls, none cut at the cap; about 58 games an hour at 27 in flight. Box spend $11.36 at list.

    calls=669 · games_finished=102 · calls_cut_at_cap=0 · box_spend_usd_list=11.36

    reported · received · @stw2 via agent · attempt only

  4. Progress report

    #4

    04:48Z: 294/460 games finished (T1 30, T2 60, T3 150, T4 52, T5 2), 2,076 calls, none cut at the cap; about 55 games an hour. Box spend $32.08 at list.

    calls=2076 · games_finished=294 · calls_cut_at_cap=0 · box_spend_usd_list=32.08

    reported · received · @stw2 via agent · attempt only

  5. Succeeded

    #5

    Exit 0; 460/460 dev games finished, score --verify verified every game. T2-T6 headline 10.4 [7.8, 15.1]; seed-only 4.7; EIG 0.70 bits per experiment, 5.4 experiments per game. Run files restricted (sha256 given); summaries public at f32ffe2. Format: 7.7% of headline decisions malformed, nearly all submissions whose rule fails the schema (243 of 244 invalid turns, kind schema_error); no call reached the cap. Box spend about $52 at list over 9 h.

    Exit code 0. Duration 31745 s.

    T1.wins=14 · T2.wins=15 · T3.wins=30 · T4.wins=0 · T5.wins=7 · T6.wins=0 · T1.scored=30 · T2.scored=60 · …

    • qwen3.5-27b-fp8-vllm-h100.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.5-27b-fp8-vllm-h100.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 e61343629098… · 133831045 bytes · M21
    • qwen3.5-27b-fp8-vllm-h100.score.json · https://github.com/stw2/zendo-lab/blob/f32ffe27ebec438161a8d9c2257be03fa19a38a5/experiments/E01-dev-baseline/results/scores/qwen3.5-27b-fp8-vllm-h100.score.json · public · sha256 18eae046911b… · 24695 bytes · M22
    • qwen3.5-27b-fp8-vllm-h100.diagnose.md · https://github.com/stw2/zendo-lab/blob/f32ffe27ebec438161a8d9c2257be03fa19a38a5/experiments/E01-dev-baseline/results/scores/qwen3.5-27b-fp8-vllm-h100.diagnose.md · public · sha256 cee137a2701e… · 1620 bytes · M23
    • qwen3.5-27b-fp8-vllm-h100.files.json · https://github.com/stw2/zendo-lab/blob/f32ffe27ebec438161a8d9c2257be03fa19a38a5/experiments/E01-dev-baseline/results/scores/qwen3.5-27b-fp8-vllm-h100.files.json · public · sha256 947fe6673c3e… · 156 bytes · M24

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.