Finding

Sign in with GitHub
← Publications

Finding · P242 · Author-curated

On the tool-use tasks with at most 8 steps and 2048 generated tokens per step, 0 of the 3 task families have a 95% bootstrap interval of the end-to-end success gap between Bielik-11B-v3.0-Instruct and Bielik-PL-11B-v3.0-Instruct that excludes 0.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Intervals excluding zero

metric
End-to-end success rateQuantity whose difference is taken.
benchmark
Agentic task suiteTasks measured.
setting
ReAct tool loop with a persistent Python sandbox: at most 8 steps, 2048 generated tokens per step, a FINAL answer also ends a trajectory in a turn with code once earlier code has executed; write-and-debug solutions rebuilt from executed code when no solution.py was written; MLX bf16, greedy decodingConditions of the measurement.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
groups
CSV analysis; methods-paragraph reproduction; write-and-debugGroups of cells compared.
count excluding zero
0 familiesGroups whose interval excludes 0.
count total
3 familiesGroups compared.

Experimental provenance

Method and evaluation protocol
Bootstrap the paired success difference within each task family (2000 resamples, seed 20260707).
Dataset
100 synthetic task instances from 20 parametric templates, each with English and Polish instructions.Version: tasks manifest sha256 66029e45c11f07613d698a67498b2c2c29faee4f0acc30bcf667bc3af6306f50 · Access: public
Reported results
CSV analysis 0.0375 [-0.0125, 0.0875] over 80 cells; methods-paragraph reproduction -0.05 [-0.2, 0.1] over 60 cells; write-and-debug 0.0833 [-0.1167, 0.2833] over 60 cells.
Uncertainty and replication
95% percentile bootstrap intervals, 2000 resamples each; no multiplicity adjustment.
Limitations
Single greedy run per cell on synthetic tasks; the two models differ by tokenizer, continued pretraining and post-training at once.

Author’s note

Values from results/analysis.json (delta_ee.csv, delta_ee.methods, delta_ee.debug) at tag e10-run2-results.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Intervals excluding zero

Of count_total groups of cells, count_excluding_zero have a 95% paired bootstrap interval of the comparator-minus-subject difference of the metric that excludes 0.

Key intervals_excluding_zero · version c5d56bcd-791d-482d-984c-21237d0ed7ad

Agentic task suite

100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

Key agentic_task_suite · version 0ea2c3bc-bc27-4395-9371-ce71dbf2c962

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

End-to-end success rate

Share of task cells a model completes with a correct FINAL value or a solution that passes the hidden tests; trajectories that end without an accepted FINAL answer count as failures.

Key end_to_end_success_rate · version 8aa03375-42ba-4d69-9f95-cb05758fbfcf

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04