Execution attempt · A3 · planned
E1 arm gpt-oss-120b-openrouter-high: all 460 dev games, run per the accepted plan
Pinned source and configuration
https://github.com/stw2/zendo-lab @ fac2c1b7381bf8c7a821dff4c32c55fbb2b290b0
Reference checked 2026-10-04 23:20 UTC. Later commits, branches or plan changes do not retarget this attempt.
- experiments/E01-dev-baseline/DESIGN.md
- experiments/E01-dev-baseline/scripts/01_run.sh
- experiments/E01-dev-baseline/scripts/02_score.sh
- experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json
- experiments/E01-dev-baseline/scripts/daytona/ctl.py
- experiments/E01-dev-baseline/scripts/daytona/guard.py
- experiments/E01-dev-baseline/scripts/daytona/setup.sh
- experiments/E01-dev-baseline/scripts/daytona/serve.sh
- pyproject.toml
- uv.lock
- Command
- bash experiments/E01-dev-baseline/scripts/01_run.sh gpt-oss-120b-openrouter-high
- Working directory
- .
- Configuration paths
- experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json experiments/E01-dev-baseline/scripts/daytona/serve.sh
- Parameters
- model=openai/gpt-oss-120b; backend=OpenRouter, default routing, 64 calls in flight; hardware=provider; sampling=reasoning effort high; max_tokens=65536
- Environment
- Harness on macOS (Python 3.13, ZendoBench 1.0.0) over HTTPS to OpenRouter
- Output directory
- experiments/E01-dev-baseline/results/runs
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- zendo-bench-1.0.0-dev-manifest · evaluation data · M1 zendo-bench-1.0.0-dev-manifest · dataset · Download
Latest reported metrics
Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.
- T1.wins
- 23
- T2.wins
- 19
- T3.wins
- 43
- T4.wins
- 0
- T5.wins
- 5
- T6.wins
- 3
- T1.scored
- 30
- T2.scored
- 60
- T3.scored
- 150
- T4.scored
- 60
- T5.scored
- 100
- T6.scored
- 60
- T1.win_rate
- 0.7666666666666667
- T2.win_rate
- 0.31666666666666665
- T3.win_rate
- 0.2866666666666667
- T4.win_rate
- 0
- T5.win_rate
- 0.05
- T6.win_rate
- 0.05
- alive_first
- 12
- diagnose.eig
- 0.6480110533174602
- diagnose.zero
- 0.18018898664059954
- headline_items
- 430
- headline_t2_t6
- 0.14066666666666666
- malformed_rate
- 0.01866378144441439
and 14 more in the attempt JSON.
Delivered events
Registered
#1E1 arm gpt-oss-120b-openrouter-high: all 460 dev games, run per the accepted plan
Started
#2Launched 2026-10-04T23:3x Z at fac2c1b (clean tree): bash experiments/E01-dev-baseline/scripts/01_run.sh gpt-oss-120b-openrouter-high. The first launch was attached to a session shell and was stopped after about 3 minutes (SIGTERM, exit 143, finished games kept in gpt-oss-120b-openrouter-high.jsonl); it resumed detached into gpt-oss-120b-openrouter-high.part2.jsonl with --exclude-finished, as the plan provides.
Succeeded
#3Exit 0; 460/460 dev games finished, score --verify verified every game. T2-T6 headline 14.1 [11.1, 19.2]; seed-only 4.7; EIG 0.65 bits per experiment, 6.7 experiments per game. Run files restricted (sha256 given); summaries public at 6b68467. Provider routing: OpenRouter's default routing served the run from several providers; the harness logged 241 system_fingerprint changes (calls by fingerprint: none reported 3465, vllm-0.27.1-tp4 336, fastcoe 73, vllm-v0.22.0-tp4 40, OpenAI fp_* 22). Exit codes by run file: [143, 0] (an earlier part was stopped and resumed).
Exit code 0. Duration 6648 s.
- gpt-oss-120b-openrouter-high.jsonl · experiments/E01-dev-baseline/results/runs/gpt-oss-120b-openrouter-high.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 f682ea74b17a… · 74689 bytes · M7
- gpt-oss-120b-openrouter-high.part2.jsonl · experiments/E01-dev-baseline/results/runs/gpt-oss-120b-openrouter-high.part2.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 b0654289a871… · 114621779 bytes · M8
- gpt-oss-120b-openrouter-high.score.json · https://github.com/stw2/zendo-lab/blob/6b684670d6023ba6d309939ac3139e4f98467276/experiments/E01-dev-baseline/results/scores/gpt-oss-120b-openrouter-high.score.json · public · sha256 46c854cf12fe… · 25681 bytes · M9
- gpt-oss-120b-openrouter-high.diagnose.md · https://github.com/stw2/zendo-lab/blob/6b684670d6023ba6d309939ac3139e4f98467276/experiments/E01-dev-baseline/results/scores/gpt-oss-120b-openrouter-high.diagnose.md · public · sha256 df63e06a67c4… · 1643 bytes · M10
- gpt-oss-120b-openrouter-high.files.json · https://github.com/stw2/zendo-lab/blob/6b684670d6023ba6d309939ac3139e4f98467276/experiments/E01-dev-baseline/results/scores/gpt-oss-120b-openrouter-high.files.json · public · sha256 ce0a5513c193… · 317 bytes · M11
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.