Finding

Sign in with GitHub
← Publications

Finding · P239 · Author-curated

On the tool-use tasks with at most 8 steps and 2048 generated tokens per step, the end-to-end success gap of Bielik-11B-v3.0-Instruct over Bielik-PL-11B-v3.0-Instruct is -0.02 with English instructions and 0.07 with Polish instructions, a difference of -0.09 (95% bootstrap interval -0.24 to 0.07).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Difference in differences

metric
End-to-end success rateQuantity whose difference is taken.
benchmark
Agentic task suiteTasks measured.
setting
ReAct tool loop with a persistent Python sandbox: at most 8 steps, 2048 generated tokens per step, a FINAL answer also ends a trajectory in a turn with code once earlier code has executed; write-and-debug solutions rebuilt from executed code when no solution.py was written; MLX bf16, greedy decodingConditions of the measurement.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
english difference
-0.02 proportionComparator minus subject on English-instructed cells.
polish difference
0.07 proportionComparator minus subject on Polish-instructed cells.
value
-0.09 proportionEnglish difference minus Polish difference.
interval low
-0.24 proportionLower bound of the 95% bootstrap interval.
interval high
0.07 proportionUpper bound of the 95% bootstrap interval.

Experimental provenance

Method and evaluation protocol
Take the paired success difference on the 100 English-instructed and the 100 Polish-instructed cells, subtract, and bootstrap the difference (2000 resamples, seed 20260707); Holm adjustment over H2 and H3.
Dataset
100 synthetic task instances from 20 parametric templates, each with English and Polish instructions.Version: tasks manifest sha256 66029e45c11f07613d698a67498b2c2c29faee4f0acc30bcf667bc3af6306f50 · Access: public
Reported results
DD -0.09 [-0.24, 0.07]; bootstrap p 0.291, Holm-adjusted 0.328.
Uncertainty and replication
95% percentile bootstrap interval, 2000 resamples; p is twice the smaller share of resamples on either side of 0.
Limitations
The bootstrap resamples English and Polish cells independently, not by instance. Single greedy run per cell on synthetic tasks; the two models differ by tokenizer, continued pretraining and post-training at once.

Author’s note

Values from results/analysis.json (H2, holm_p, delta_ee) at tag e10-run2-results.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Difference in differences

The comparator-minus-subject difference of the metric is english_difference on English-instructed cells and polish_difference on Polish-instructed cells; value is english_difference minus polish_difference, with a 95% bootstrap interval from interval_low to interval_high.

Key difference_in_differences · version 4d4c5ae8-f9c8-4588-905e-c88e4f7f7ac6

Agentic task suite

100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

Key agentic_task_suite · version 0ea2c3bc-bc27-4395-9371-ce71dbf2c962

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

End-to-end success rate

Share of task cells a model completes with a correct FINAL value or a solution that passes the hidden tests; trajectories that end without an accepted FINAL answer count as failures.

Key end_to_end_success_rate · version 8aa03375-42ba-4d69-9f95-cb05758fbfcf

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04