Execution attempt · A6 · planned
E1 arm qwen3.5-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan
Pinned source and configuration
https://github.com/stw2/zendo-lab @ fac2c1b7381bf8c7a821dff4c32c55fbb2b290b0
Reference checked 2026-10-04 23:27 UTC. Later commits, branches or plan changes do not retarget this attempt.
- experiments/E01-dev-baseline/DESIGN.md
- experiments/E01-dev-baseline/scripts/01_run.sh
- experiments/E01-dev-baseline/scripts/02_score.sh
- experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json
- experiments/E01-dev-baseline/scripts/daytona/ctl.py
- experiments/E01-dev-baseline/scripts/daytona/guard.py
- experiments/E01-dev-baseline/scripts/daytona/setup.sh
- experiments/E01-dev-baseline/scripts/daytona/serve.sh
- pyproject.toml
- uv.lock
- Command
- bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-27b-fp8-vllm-h100
- Working directory
- .
- Configuration paths
- experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json experiments/E01-dev-baseline/scripts/daytona/serve.sh
- Parameters
- model=Qwen/Qwen3.5-27B-FP8@97f5941b; backend=vLLM 0.30.0, --batch auto; hardware=1x H100 80GB (Daytona, on-demand); sampling=temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 1.5; max_tokens=65536
- Environment
- Harness on macOS (Python 3.13, ZendoBench 1.0.0) over HTTPS to vLLM 0.30.0 on one on-demand H100 80GB (Daytona sandbox, CUDA 13.0 image)
- Output directory
- experiments/E01-dev-baseline/results/runs
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · dataset · Download
Latest reported metrics
Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.
- T1.wins
- 14
- T2.wins
- 15
- T3.wins
- 30
- T4.wins
- 0
- T5.wins
- 7
- T6.wins
- 0
- T1.scored
- 30
- T2.scored
- 60
- T3.scored
- 150
- T4.scored
- 60
- T5.scored
- 100
- T6.scored
- 60
- T1.win_rate
- 0.4666666666666667
- T2.win_rate
- 0.25
- T3.win_rate
- 0.2
- T4.win_rate
- 0
- T5.win_rate
- 0.07
- T6.win_rate
- 0
- alive_first
- 27
- diagnose.eig
- 0.6970935985259261
- diagnose.zero
- 0.12738336713995943
- headline_items
- 430
- headline_t2_t6
- 0.10400000000000001
- malformed_rate
- 0.0767054385413392
and 14 more in the attempt JSON.
Delivered events
Registered
#1E1 arm qwen3.5-27b-fp8-vllm-h100: all 460 dev games, run per the accepted plan
Started
#2Launched 2026-10-04T23:28:05Z at fac2c1b (clean tree): bash experiments/E01-dev-baseline/scripts/01_run.sh qwen3.5-27b-fp8-vllm-h100, against vLLM 0.30.0 on one on-demand H100 (model files matched the Hub's hashes; chat check passed); 27 games in flight (--batch auto). Guard: stops at $200 or 34 h.
Progress report
#301:14Z: 102/460 games finished (T1 30, T2 60, T3 12), 669 calls, none cut at the cap; about 58 games an hour at 27 in flight. Box spend $11.36 at list.
Progress report
#404:48Z: 294/460 games finished (T1 30, T2 60, T3 150, T4 52, T5 2), 2,076 calls, none cut at the cap; about 55 games an hour. Box spend $32.08 at list.
Succeeded
#5Exit 0; 460/460 dev games finished, score --verify verified every game. T2-T6 headline 10.4 [7.8, 15.1]; seed-only 4.7; EIG 0.70 bits per experiment, 5.4 experiments per game. Run files restricted (sha256 given); summaries public at f32ffe2. Format: 7.7% of headline decisions malformed, nearly all submissions whose rule fails the schema (243 of 244 invalid turns, kind schema_error); no call reached the cap. Box spend about $52 at list over 9 h.
Exit code 0. Duration 31745 s.
- qwen3.5-27b-fp8-vllm-h100.jsonl · experiments/E01-dev-baseline/results/runs/qwen3.5-27b-fp8-vllm-h100.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 e61343629098… · 133831045 bytes · M21
- qwen3.5-27b-fp8-vllm-h100.score.json · https://github.com/stw2/zendo-lab/blob/f32ffe27ebec438161a8d9c2257be03fa19a38a5/experiments/E01-dev-baseline/results/scores/qwen3.5-27b-fp8-vllm-h100.score.json · public · sha256 18eae046911b… · 24695 bytes · M22
- qwen3.5-27b-fp8-vllm-h100.diagnose.md · https://github.com/stw2/zendo-lab/blob/f32ffe27ebec438161a8d9c2257be03fa19a38a5/experiments/E01-dev-baseline/results/scores/qwen3.5-27b-fp8-vllm-h100.diagnose.md · public · sha256 cee137a2701e… · 1620 bytes · M23
- qwen3.5-27b-fp8-vllm-h100.files.json · https://github.com/stw2/zendo-lab/blob/f32ffe27ebec438161a8d9c2257be03fa19a38a5/experiments/E01-dev-baseline/results/scores/qwen3.5-27b-fp8-vllm-h100.files.json · public · sha256 947fe6673c3e… · 156 bytes · M24
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.