Finding

Sign in with GitHub
← Publications

Finding · P238 · Author-curated

On 200 paired tool-use task cells with English and Polish instructions, at most 8 steps and 2048 generated tokens per step, Bielik-PL-11B-v3.0-Instruct succeeds end to end on 0.7 of cells and Bielik-11B-v3.0-Instruct on 0.725, a gap of 0.025 (95% bootstrap interval -0.055 to 0.105).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric difference with an interval

metric
End-to-end success rateQuantity compared.
benchmark
Agentic task suiteTasks measured.
language
English and Polish instructionsInstruction languages of the cells.
setting
ReAct tool loop with a persistent Python sandbox: at most 8 steps, 2048 generated tokens per step, a FINAL answer also ends a trajectory in a turn with code once earlier code has executed; write-and-debug solutions rebuilt from executed code when no solution.py was written; MLX bf16, greedy decodingConditions of the measurement.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
subject value
0.7 proportionSuccess rate of the subject.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
comparator value
0.725 proportionSuccess rate of the comparator.
difference
0.025 proportionComparator value minus subject value.
interval low
-0.055 proportionLower bound of the 95% bootstrap interval.
interval high
0.105 proportionUpper bound of the 95% bootstrap interval.
cells
200 task cellsPaired task cells.

Experimental provenance

Method and evaluation protocol
Grade each trajectory (numeric tolerance or hidden tests), count trajectories without an accepted FINAL answer as failures, and bootstrap the mean paired difference over the cells (2000 resamples, seed 20260707).
Dataset
100 synthetic task instances from 20 parametric templates, each with English and Polish instructions.Version: tasks manifest sha256 66029e45c11f07613d698a67498b2c2c29faee4f0acc30bcf667bc3af6306f50 · Access: public
Reported results
Success 0.725 (original) and 0.7 (APT4 model); gap 0.025 [-0.055, 0.105]; English instructions -0.02 [-0.12, 0.07]; Polish instructions 0.07 [-0.06, 0.19].
Uncertainty and replication
95% percentile bootstrap interval over paired cells, 2000 resamples.
Limitations
Single greedy run per cell on synthetic tasks; the two models differ by tokenizer, continued pretraining and post-training at once.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Values from results/analysis.json (delta_ee) and results/graded_summary.json (overall) at tag e10-run2-results; flat in results/metrics.json.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric difference with an interval

The subject and the comparator take subject_value and comparator_value of the metric on the same cells; difference is comparator_value minus subject_value, with a 95% paired bootstrap interval from interval_low to interval_high over cells task cells; benchmark, language and setting state what was measured.

Key metric_difference_with_interval · version 8aa03375-42ba-4d69-9f95-cb05758fbfcf

Agentic task suite

100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

Key agentic_task_suite · version 0ea2c3bc-bc27-4395-9371-ce71dbf2c962

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

End-to-end success rate

Share of task cells a model completes with a correct FINAL value or a solution that passes the hidden tests; trajectories that end without an accepted FINAL answer count as failures.

Key end_to_end_success_rate · version 8aa03375-42ba-4d69-9f95-cb05758fbfcf

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

extends

15132b6d-26e7-49b9-8217-b144a28236c7

Compares the pair on multi-step tool-use tasks beyond the nine benchmarks.