Finding

Sign in with GitHub
← Publications

Finding · P240 · Author-curated

On the 100 English-instructed tool-use task cells with at most 8 steps and 2048 generated tokens per step, Bielik-PL-11B-v3.0-Instruct generates 806.6 tokens per trajectory on average and Bielik-11B-v3.0-Instruct 1011.6, a ratio of 0.7974 (95% bootstrap interval 0.5017 to 1.0846).

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Metric ratio with an interval

metric
Mean generated tokens per trajectoryQuantity compared.
benchmark
Agentic task suiteTasks measured.
language
EnglishInstruction language of the cells.
setting
ReAct tool loop with a persistent Python sandbox: at most 8 steps, 2048 generated tokens per step, a FINAL answer also ends a trajectory in a turn with code once earlier code has executed; write-and-debug solutions rebuilt from executed code when no solution.py was written; MLX bf16, greedy decodingConditions of the measurement.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
subject value
806.6 tokensMean for the subject.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
comparator value
1011.6 tokensMean for the comparator.
value
0.7974 ratioSubject mean divided by comparator mean.
interval low
0.5017 ratioLower bound of the 95% bootstrap interval.
interval high
1.0846 ratioUpper bound of the 95% bootstrap interval.

Experimental provenance

Method and evaluation protocol
Divide the mean generated tokens per English-instructed trajectory of the APT4 model by the original's and bootstrap the ratio over paired cells (2000 resamples, seed 20260707).
Dataset
100 synthetic task instances from 20 parametric templates, each with English and Polish instructions.Version: tasks manifest sha256 66029e45c11f07613d698a67498b2c2c29faee4f0acc30bcf667bc3af6306f50 · Access: public
Reported results
Ratio 0.7974 [0.5017, 1.0846]; bootstrap p for a ratio of 1: 0.164, Holm-adjusted 0.328; the original model's turns contain a <think> block in 55 of 100 English cells.
Uncertainty and replication
95% percentile bootstrap interval over paired cells, 2000 resamples.
Limitations
H3 as designed compares prompt-side tokens; the analysis computes generated tokens. Single greedy run per cell on synthetic tasks; the two models differ by tokenizer, continued pretraining and post-training at once.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author

Associated through an explicitly referenced cited claim.

Author’s note

Values from results/analysis.json (H3) and results/graded_summary.json (by_lang.en) at tag e10-run2-results.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Metric ratio with an interval

The subject and the comparator take subject_value and comparator_value of the metric on the same cells; value is subject_value divided by comparator_value, with a 95% paired bootstrap interval from interval_low to interval_high.

Key metric_ratio_with_interval · version 0ea16ba0-68e0-49a0-934c-ca8946e78a80

Agentic task suite

100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

Key agentic_task_suite · version 0ea2c3bc-bc27-4395-9371-ce71dbf2c962

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Mean generated tokens per trajectory

Mean over trajectories of the tokens a model generates, summed over the trajectory's steps.

Key mean_generated_tokens · version 0ea16ba0-68e0-49a0-934c-ca8946e78a80

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

related

eed1981a-09e9-4aa9-bd57-3db32c7ba4a0

English fertility of the two tokenizers.

related

64285728-7520-4c5e-9d7a-9dbbe2c874f7

The original model's <think> tag on English retrieval prompts.