Execution attempt · A5 · planned
E1 arm qwen3.8-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan
Pinned source and configuration
https://github.com/stw2/zendo-lab @ fac2c1b7381bf8c7a821dff4c32c55fbb2b290b0
Reference checked 2026-10-04 23:27 UTC. Later commits, branches or plan changes do not retarget this attempt.
- experiments/E01-dev-baseline/DESIGN.md
- experiments/E01-dev-baseline/scripts/01_run.sh
- experiments/E01-dev-baseline/scripts/02_score.sh
- experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json
- experiments/E01-dev-baseline/scripts/daytona/ctl.py
- experiments/E01-dev-baseline/scripts/daytona/guard.py
- experiments/E01-dev-baseline/scripts/daytona/setup.sh
- experiments/E01-dev-baseline/scripts/daytona/serve.sh
- pyproject.toml
- uv.lock
- Command
- bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.8-27b-fp8-vllm-h100
- Working directory
- .
- Configuration paths
- experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json experiments/E01-dev-baseline/scripts/daytona/serve.sh
- Parameters
- model=Qwen/Qwen3.8-27B-FP8@017b9c7a; backend=vLLM 0.30.0, --batch auto; hardware=1x H100 80GB (Daytona, on-demand); sampling=temperature 1.0, top_p 0.95, top_k 20, presence_penalty 0, reasoning_effort xhigh; max_tokens=65536
- Environment
- Harness on macOS (Python 3.13, ZendoBench 1.0.0) over HTTPS to vLLM 0.30.0 on one on-demand H100 80GB (Daytona sandbox, CUDA 13.0 image)
- Output directory
- experiments/E01-dev-baseline/results/runs
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · dataset · Download
Latest reported metrics
Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.
- T1.wins
- 29
- T2.wins
- 50
- T3.wins
- 118
- T4.wins
- 5
- T5.wins
- 34
- T6.wins
- 11
- T1.scored
- 30
- T2.scored
- 60
- T3.scored
- 150
- T4.scored
- 60
- T5.scored
- 100
- T6.scored
- 60
- T1.win_rate
- 0.9666666666666667
- T2.win_rate
- 0.8333333333333334
- T3.win_rate
- 0.7866666666666666
- T4.win_rate
- 0.08333333333333334
- T5.win_rate
- 0.33999999999999997
- T6.win_rate
- 0.18333333333333332
- alive_first
- 2
- diagnose.eig
- 0.44566157924183153
- diagnose.zero
- 0.4538291837377876
- headline_items
- 430
- headline_t2_t6
- 0.44533333333333336
- malformed_rate
- 0.007673011657075402
and 14 more in the attempt JSON.
Delivered events
Registered
#1E1 arm qwen3.8-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan
Started
#2Launched 2026-10-04T23:27:22Z at fac2c1b (clean tree): bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.8-27b-fp8-vllm-h100, against vLLM 0.30.0 on one on-demand H100 (model files matched the Hub's hashes; chat check passed). Guard: stops at $200 or 34 h.
Progress report
#301:14Z: 34/460 games finished (T1 29, T2 5), 394 calls, none cut at the cap; about 19 games an hour at 27 in flight. Box spend $11.29 at list.
Progress report
#404:48Z: 107/460 games finished (T1 30, T2 60, T3 17), 1,464 calls, none cut at the cap; about 20 games an hour. Box spend $32.11 at list.
Progress report
#5Interrupted and resuming. At 05:58Z the first H100 box stopped answering: its Daytona preview returned HTTP 502 (5 games exhausted their retries), then timed out, and the harness's breaker stopped the run (exit 4). 129 games had finished; 31 are unfinished and will be replayed. The box was deleted at 06:03Z ($39.33 at list to then). A replacement box (same image, model revision 017b9c7a, vLLM 0.30.0, server flags) launched at 07:27Z; the run resumes into qwen3.8-27b-fp8-vllm-h100.part2.jsonl with --exclude-finished as this same attempt. Amendment in DESIGN.md at e657b11.
Progress report
#6Interrupted again. The guard's Daytona login expired at 09:36Z, so it could not renew the second box's server-side TTL; the box was destroyed at about 12:31Z (unverified until the next login) and the harness's breaker stopped part2. Finished so far: 223 of 460 games (129 in the first file, 94 in part2); 29 unfinished games of part2 are replayed. Box spend so far $68.69 at list for both boxes. The run resumes on a third box into part3 with --exclude-finished, as this same attempt, once the Daytona login is renewed; the guard now rebuilds its client after a login error.
Progress report
#7Interrupted a second time and resuming. The guard's Daytona credential (the CLI's 24-hour login) expired at 09:36Z, so the second box's server-side TTL was not renewed and Daytona destroyed it at about 12:30Z; the run's breaker stopped part 2. 223/460 games finished across parts 1 and 2 (part 2: 94 finished). Box spend so far $68.69 at list for two boxes. A third box (same image, model revision, vLLM and flags) launched at 14:40Z; the run resumes into qwen3.8-27b-fp8-vllm-h100.part3.jsonl with --exclude-finished, as this same attempt. The tooling now uses a non-expiring Daytona API key; amendment at 38a396e.
Progress report
#819:39Z: 283/460 games finished (T1-T3 complete, T4 43/60), part 3 running without stops. Box spend $97.62 at list across three boxes; the remaining T4-T6 games project to about $96 more, so the owner raised the cost cap from $200 to $240 (amendment at bee32e4).
Succeeded
#9Exit 0; 460/460 dev games finished, score --verify verified every game. T2-T6 headline 44.5 [40.4, 49.0]; seed-only 4.7; EIG 0.45 bits per experiment, 13.8 experiments per game. Run files restricted (sha256 given); summaries public at ed6864b. Run across three H100 boxes (DESIGN.md amendments): box 1 hung at 05:58Z and box 2 was destroyed when its TTL lapsed at about 12:30Z; each time the harness's breaker stopped the run and it resumed with --exclude-finished, and the 60 games unfinished at those stops were replayed. Boxes $168.46 at list in total. Exit codes by run file: [4, 0] (an earlier part was stopped and resumed).
Exit code 0. Duration 115491 s.
- qwen3.8-27b-fp8-vllm-h100.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 3a1ee9069562… · 104181802 bytes · M29
- qwen3.8-27b-fp8-vllm-h100.part2.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.part2.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 93642eb59f44… · 78559599 bytes · M30
- qwen3.8-27b-fp8-vllm-h100.part3.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.8-27b-fp8-vllm-h100.part3.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 97bbeb22d449… · 242721721 bytes · M31
- qwen3.8-27b-fp8-vllm-h100.score.json · https://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.score.json · public · sha256 9cee7ea1aae3… · 27189 bytes · M32
- qwen3.8-27b-fp8-vllm-h100.diagnose.md · https://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.diagnose.md · public · sha256 470ab62ce1ac… · 1628 bytes · M33
- qwen3.8-27b-fp8-vllm-h100.files.json · https://github.com/stw2/zendo-lab/blob/ed6864b84e9d143430c580bfec42875b2898b49e/experiments/E01-dev-baseline/results/scores/qwen3.8-27b-fp8-vllm-h100.files.json · public · sha256 e5c873c56950… · 473 bytes · M34
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.