Execution attempt · A96 · planned
Poses each instance once, without the tool loop, to Bielik-11B-v3.0-Instruct as a per-step probe in both languages, and aggregates both models' probe success rates.
Pinned source and configuration
https://github.com/stw2/tokenizer-science-tax @ 03e30b6cecb9c899914b18931b1f8ee0bdc7d2f5
Reference checked 2026-09-14 13:17 UTC. Later commits, branches or plan changes do not retarget this attempt.
- Command
- python scripts/02_run_static.py --model orig
- Working directory
- experiments/E10-agentic-compounding
- Configuration paths
- None
- Parameters
- model orig; languages en,pl; one shot; 900 generated tokens; greedy
- Environment
- 3.12.13 (uv 0.11.16); MLX bf16 on Metal, greedy decoding (temperature 0); Apple silicon, 128 GB unified memory
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- tasks_manifest.json · configuration · M457 Agentic task manifest · configuration · Download
- data/instances · evaluation data · M459 Agentic task instances · dataset · Ask the reporter
- orig MLX bf16 conversion · checkpoint · M66 bielik-11b-v3.0-instruct mlx-bf16 · raw output · Ask the reporter
Delivered events
Registered
#1Poses each instance once, without the tool loop, to Bielik-11B-v3.0-Instruct as a per-step probe in both languages, and aggregates both models' probe success rates.
Started
#2Started.
Succeeded
#3Completed with exit code 0.
Exit code 0.
- static_orig.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/7458cbb2d1925fda6196c42e14dca7ccbc3b295c/experiments/E10-agentic-compounding/results/raw/static_orig.jsonl · public · sha256 20d6dd2d7a9a… · 24276 bytes · M465
- static_probs.json · https://github.com/stw2/tokenizer-science-tax/blob/7458cbb2d1925fda6196c42e14dca7ccbc3b295c/experiments/E10-agentic-compounding/results/static_probs.json · public · sha256 002e33febd8c… · 365 bytes · M466
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.