Execution attempt

Sign in with GitHub
← Experiment E1 · Where do small and strong models stand on ZendoBench 1.0.0 dev, and how do they play: how much information do their experiments gain, and how many rule classes are still alive when they submit?

Execution attempt · A3 · planned

E1 arm gpt-oss-120b-openrouter-high: all 460 dev games, run per the accepted plan

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/zendo-lab @ fac2c1b7381bf8c7a821dff4c32c55fbb2b290b0

Reference checked 2026-10-04 23:20 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
bash experiments/E01-dev-baseline/scripts/01_run.sh gpt-oss-120b-openrouter-high
Working directory
.
Configuration paths
experiments/E01-dev-baseline/scripts/checkpoints/qwen3.5-2b.json experiments/E01-dev-baseline/scripts/daytona/serve.sh
Parameters
model=openai/gpt-oss-120b; backend=OpenRouter, default routing, 64 calls in flight; hardware=provider; sampling=reasoning effort high; max_tokens=65536
Environment
Harness on macOS (Python 3.13, ZendoBench 1.0.0) over HTTPS to OpenRouter
Output directory
experiments/E01-dev-baseline/results/runs

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

Latest reported metrics

Values the registrant’s tool attached to its latest progress or outcome event. Substrate stores them; it verifies no measurement.

T1.wins
23
T2.wins
19
T3.wins
43
T4.wins
0
T5.wins
5
T6.wins
3
T1.scored
30
T2.scored
60
T3.scored
150
T4.scored
60
T5.scored
100
T6.scored
60
T1.win_rate
0.7666666666666667
T2.win_rate
0.31666666666666665
T3.win_rate
0.2866666666666667
T4.win_rate
0
T5.win_rate
0.05
T6.win_rate
0.05
alive_first
12
diagnose.eig
0.6480110533174602
diagnose.zero
0.18018898664059954
headline_items
430
headline_t2_t6
0.14066666666666666
malformed_rate
0.01866378144441439

and 14 more in the attempt JSON.

Delivered events

  1. Registered

    #1

    E1 arm gpt-oss-120b-openrouter-high: all 460 dev games, run per the accepted plan

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Launched 2026-10-04T23:3x Z at fac2c1b (clean tree): bash experiments/E01-dev-baseline/scripts/01_run.sh gpt-oss-120b-openrouter-high. The first launch was attached to a session shell and was stopped after about 3 minutes (SIGTERM, exit 143, finished games kept in gpt-oss-120b-openrouter-high.jsonl); it resumed detached into gpt-oss-120b-openrouter-high.part2.jsonl with --exclude-finished, as the plan provides.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Exit 0; 460/460 dev games finished, score --verify verified every game. T2-T6 headline 14.1 [11.1, 19.2]; seed-only 4.7; EIG 0.65 bits per experiment, 6.7 experiments per game. Run files restricted (sha256 given); summaries public at 6b68467. Provider routing: OpenRouter's default routing served the run from several providers; the harness logged 241 system_fingerprint changes (calls by fingerprint: none reported 3465, vllm-0.27.1-tp4 336, fastcoe 73, vllm-v0.22.0-tp4 40, OpenAI fp_* 22). Exit codes by run file: [143, 0] (an earlier part was stopped and resumed).

    Exit code 0. Duration 6648 s.

    T1.wins=23 · T2.wins=19 · T3.wins=43 · T4.wins=0 · T5.wins=5 · T6.wins=3 · T1.scored=30 · T2.scored=60 · …

    • gpt-oss-120b-openrouter-high.jsonl · experiments/E01-dev-baseline/results/runs/gpt-oss-120b-openrouter-high.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 f682ea74b17a… · 74689 bytes · M7
    • gpt-oss-120b-openrouter-high.part2.jsonl · experiments/E01-dev-baseline/results/runs/gpt-oss-120b-openrouter-high.part2.jsonl (held by the Room owner; not public; ask the reporter) · restricted · sha256 b0654289a871… · 114621779 bytes · M8
    • gpt-oss-120b-openrouter-high.score.json · https://github.com/stw2/zendo-lab/blob/6b684670d6023ba6d309939ac3139e4f98467276/experiments/E01-dev-baseline/results/scores/gpt-oss-120b-openrouter-high.score.json · public · sha256 46c854cf12fe… · 25681 bytes · M9
    • gpt-oss-120b-openrouter-high.diagnose.md · https://github.com/stw2/zendo-lab/blob/6b684670d6023ba6d309939ac3139e4f98467276/experiments/E01-dev-baseline/results/scores/gpt-oss-120b-openrouter-high.diagnose.md · public · sha256 df63e06a67c4… · 1643 bytes · M10
    • gpt-oss-120b-openrouter-high.files.json · https://github.com/stw2/zendo-lab/blob/6b684670d6023ba6d309939ac3139e4f98467276/experiments/E01-dev-baseline/results/scores/gpt-oss-120b-openrouter-high.files.json · public · sha256 ce0a5513c193… · 317 bytes · M11

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.