Execution attempt

Sign in with GitHub
← Experiment E1 · Where do small and strong models stand on ZendoBench 1.0.0 dev, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?

Execution attempt · A5 · planned

E1 arm qwen3.8-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/zendo-lab @ fac2c1b7381bf8c7a821dff4c32c55fbb2b290b0

Reference checked 2026-10-04 23:27 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.8-27b-fp8-vllm-h100
Working directory
.
Configuration paths
experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json experiments/E01-dev-baseline/scripts/daytona/serve.sh
Parameters
model=Qwen/Qwen3.8-27B-FP8@017b9c7a; backend=vLLM 0.30.0, --batch auto; hardware=1x H100 80GB (Daytona, on-demand); sampling=temperature 1.0, top_p 0.95, top_k 20, presence_penalty 0, reasoning_effort xhigh; max_tokens=65536
Environment
Harness on macOS (Python 3.13, ZendoBench 1.0.0) over HTTPS to vLLM 0.30.0 on one on-demand H100 80GB (Daytona sandbox, CUDA 13.0 image)
Output directory
experiments/E01-dev-baseline/results/runs

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

T1.wins
29
T2.wins
50
T3.wins
118
T4.wins
5
T5.wins
34
T6.wins
11
T1.scored
30
T2.scored
60
T3.scored
150
T4.scored
60
T5.scored
100
T6.scored
60
T1.win_rate
0.9666666666666667
T2.win_rate
0.8333333333333334
T3.win_rate
0.7866666666666666
T4.win_rate
0.08333333333333334
T5.win_rate
0.33999999999999997
T6.win_rate
0.18333333333333332
alive_first
2
diagnose.eig
0.44566157924183153
diagnose.zero
0.4538291837377876
headline_items
430
headline_t2_t6
0.44533333333333336
malformed_rate
0.007673011657075402

and 14 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    E1 arm qwen3.8-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Launched 2026-10-04T23:27:22Z at fac2c1b (clean tree): bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.8-27b-fp8-vllm-h100, against vLLM 0.30.0 on one on-demand H100 (model files matched the Hub's hashes; chat check passed). Guard: stops at $200 or 34 h.

    reported · received · @stw2 via agent · attempt only

  3. Progress report

    #3

    01:14Z: 34/460 games finished (T1 29, T2 5), 394 calls, none cut at the cap; about 19 games an hour at 27 in flight. Box spend $11.29 at list.

    calls=394 · games_finished=34 · calls_cut_at_cap=0 · box_spend_usd_list=11.29

    reported · received · @stw2 via agent · attempt only

  4. Progress report

    #4

    04:48Z: 107/460 games finished (T1 30, T2 60, T3 17), 1,464 calls, none cut at the cap; about 20 games an hour. Box spend $32.11 at list.

    calls=1464 · games_finished=107 · calls_cut_at_cap=0 · box_spend_usd_list=32.11

    reported · received · @stw2 via agent · attempt only

  5. Progress report

    #5

    Interrupted and resuming. At 05:58Z the first H100 box stopped answering: its Daytona preview returned HTTP 502 (5 games exhausted their retries), then timed out, and the harness's breaker stopped the run (exit 4). 129 games had finished; 31 are unfinished and will be replayed. The box was deleted at 06:03Z ($39.33 at list to then). A replacement box (same image, model revision 017b9c7a, vLLM 0.30.0, server flags) launched at 07:27Z; the run resumes into qwen3.8-27b-fp8-vllm-h100.part2.jsonl with --exclude-finished as this same attempt. Amendment in DESIGN.md at e657b11.

    games_finished=129 · box1_spend_usd_list=39.33 · games_unfinished_replayed=31

    reported · received · @stw2 via agent · attempt only

  6. Progress report

    #6

    Interrupted again. The guard's Daytona login expired at 09:36Z, so it could not renew the second box's server-side TTL; the box was destroyed at about 12:31Z (unverified until the next login) and the harness's breaker stopped part2. Finished so far: 223 of 460 games (129 in the first file, 94 in part2); 29 unfinished games of part2 are replayed. Box spend so far $68.69 at list for both boxes. The run resumes on a third box into part3 with --exclude-finished, as this same attempt, once the Daytona login is renewed; the guard now rebuilds its client after a login error.

    games_finished=223 · box_spend_usd_list=68.69

    reported · received · @stw2 via agent · attempt only

  7. Progress report

    #7

    Interrupted a second time and resuming. The guard's Daytona credential (the CLI's 24-hour login) expired at 09:36Z, so the second box's server-side TTL was not renewed and Daytona destroyed it at about 12:30Z; the run's breaker stopped part 2. 223/460 games finished across parts 1 and 2 (part 2: 94 finished). Box spend so far $68.69 at list for two boxes. A third box (same image, model revision, vLLM and flags) launched at 14:40Z; the run resumes into qwen3.8-27b-fp8-vllm-h100.part3.jsonl with --exclude-finished, as this same attempt. The tooling now uses a non-expiring Daytona API key; amendment at 38a396e.

    games_finished=223 · games_remaining=237 · box_spend_usd_list_cumulative=68.69

    reported · received · @stw2 via agent · attempt only

  8. Progress report

    #8

    19:39Z: 283/460 games finished (T1-T3 complete, T4 43/60), part 3 running without stops. Box spend $97.62 at list across three boxes; the remaining T4-T6 games project to about $96 more, so the owner raised the cost cap from $200 to $240 (amendment at bee32e4).

    cost_cap_usd=240 · games_finished=283 · box_spend_usd_list_cumulative=97.62

    reported · received · @stw2 via agent · attempt only

  9. Succeeded

    #9

    Exit 0; 460/460 dev games finished, score --verify verified every game. T2-T6 headline 44.5 [40.4, 49.0]; seed-only 4.7; EIG 0.45 bits per experiment, 13.8 experiments per game. Run files restricted (sha256 given); summaries public at ed6864b. Run across three H100 boxes (DESIGN.md amendments): box 1 hung at 05:58Z and box 2 was destroyed when its TTL lapsed at about 12:30Z; each time the harness's breaker stopped the run and it resumed with --exclude-finished, and the 60 games unfinished at those stops were replayed. Boxes $168.46 at list in total. Exit codes by run file: [4, 0] (an earlier part was stopped and resumed).

    Exit code 0. Duration 115491 s.

    T1.wins=29 · T2.wins=50 · T3.wins=118 · T4.wins=5 · T5.wins=34 · T6.wins=11 · T1.scored=30 · T2.scored=60 · …

    • qwen3.8-27b-fp8-vllm-h100.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 3a1ee9069562… · 104181802 bytes · M29
    • qwen3.8-27b-fp8-vllm-h100.part2.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.part2.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 93642eb59f44… · 78559599 bytes · M30
    • qwen3.8-27b-fp8-vllm-h100.part3.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.part3.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 97bbeb22d449… · 242721721 bytes · M31
    • qwen3.8-27b-fp8-vllm-h100.score.json · https://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.score.json · public · sha256 9cee7ea1aae3… · 27189 bytes · M32
    • qwen3.8-27b-fp8-vllm-h100.diagnose.md · https://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.diagnose.md · public · sha256 470ab62ce1ac… · 1628 bytes · M33
    • qwen3.8-27b-fp8-vllm-h100.files.json · https://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.files.json · public · sha256 e5c873c56950… · 473 bytes · M34

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.