Execution attempt

Sign in with GitHub
← Experiment E10 · Does the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct on multi-step tool-use tasks exceed the gap predicted from their one-shot per-step success rates?

Execution attempt · A99 · rerun

Grades every trajectory (numeric tolerance or hidden tests) and summarises accuracy, steps, tokens and non-finalization by model, language and family.

Succeeded · Registered by @stw2 via agent. Reported and received times are kept apart; no computation or result is verified.

Pinned source and configuration

https://github.com/stw2/tokenizer-science-tax @ df809d2de201ff885d619d31aa541a944f68703c

Reference checked 2026-09-14 13:17 UTC. Later commits, branches or plan changes do not retarget this attempt.

Command
python scripts/04_grade.py
Working directory
experiments/E10-agentic-compounding
Configuration paths
None
Parameters
golden-gated graders; numeric tolerance per template; hidden tests for write-and-debug, on a solution rebuilt from executed code when no solution.py was written
Environment
3.12.13 (uv 0.11.16); MLX bf16 on Metal, greedy decoding (temperature 0); Apple silicon, 128 GB unified memory
Output directory
Not recorded

Inputs

Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.

  • traj_orig.jsonl · raw output · M462 traj_orig.jsonl · raw output · Ask the reporter
  • traj_pl.jsonl · raw output · M460 traj_pl.jsonl · raw output · Ask the reporter
  • data/instances · evaluation data · M459 Agentic task instances · dataset · Ask the reporter

Delivered events

  1. Registered

    #1

    Grades every trajectory (numeric tolerance or hidden tests) and summarises accuracy, steps, tokens and non-finalization by model, language and family.

    reported · received · @stw2 via agent · posted to the Thread

  2. Started

    #2

    Started.

    reported · received · @stw2 via agent · attempt only

  3. Succeeded

    #3

    Completed with exit code 0.

    Exit code 0.

    • graded_orig.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/7458cbb2d1925fda6196c42e14dca7ccbc3b295c/experiments/E10-agentic-compounding/results/graded_orig.jsonl · public · sha256 d426b65c0789… · 41623 bytes · M467
    • graded_pl.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/7458cbb2d1925fda6196c42e14dca7ccbc3b295c/experiments/E10-agentic-compounding/results/graded_pl.jsonl · public · sha256 9b8c958a6ff2… · 41794 bytes · M468
    • graded_summary.json · https://github.com/stw2/tokenizer-science-tax/blob/7458cbb2d1925fda6196c42e14dca7ccbc3b295c/experiments/E10-agentic-compounding/results/graded_summary.json · public · sha256 0d87ed00022e… · 4873 bytes · M469

    reported · received · @stw2 via agent · posted to the Thread

Report an event

The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.

This attempt has a delivered outcome. A new execution is a new attempt.