Finding

Sign in with GitHub
← Publications

Finding · P237 · Author-curated · Corrected

On the 100 English-instructed tool-use task cells with at most 8 steps and 700 generated tokens per step, 0.1 of Bielik-PL-11B-v3.0-Instruct's trajectories and 0.36 of Bielik-11B-v3.0-Instruct's end without an accepted FINAL answer.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Notices · Corrected · this exact version stays citable

Structured assertion

Relation: Metric comparison

metric
Non-finalization rateQuantity compared.
benchmark
Agentic task suiteTasks measured.
language
EnglishInstruction language of the cells.
setting
ReAct tool loop with a persistent Python sandbox: at most 8 steps, 700 generated tokens per step, a FINAL answer ends a trajectory only in a turn without code; write-and-debug solutions rebuilt from executed code when no solution.py was written; MLX bf16, greedy decodingConditions of the measurement.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
subject value
0.1 proportionRate for the subject.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
comparator value
0.36 proportionRate for the comparator.

Experimental provenance

Method and evaluation protocol
Count English-instructed trajectories that end by step-budget exhaustion, context overflow or stalling, divided by the 100 cells.
Dataset
100 synthetic task instances from 20 parametric templates, each with English and Polish instructions.Version: tasks manifest sha256 66029e45c11f07613d698a67498b2c2c29faee4f0acc30bcf667bc3af6306f50 · Access: public
Reported results
Original: 10 budget exhausted, 0 overflow, 26 stalled; APT4 model: 10 budget exhausted, 0 overflow, 0 stalled. Overflow or budget exhaustion alone: 0.1 (original) and 0.1 (APT4 model).
Uncertainty and replication
Counts over 100 cells per model; no interval.
Limitations
H3's second clause as designed counts overflow and budget exhaustion only; the analysis reports non-finalization, which also counts stalls. Single greedy run per cell on synthetic tasks; the two models differ by tokenizer, continued pretraining and post-training at once.

Author’s note

Values from results/analysis.json (H3) and results/metrics.json (outcome counts from results/graded_*.jsonl) at tag e10-run1-results.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Non-finalization rate

Share of task cells whose trajectory ends without an accepted FINAL answer: step budget used up, context window overflow, or two consecutive turns with neither code nor a FINAL answer.

Key non_finalization_rate · version 4485b760-d241-410e-949d-5a2ce6b16821

Metric comparison

The subject and the comparator take subject_value and comparator_value of the metric; benchmark, scope, evaluation_text, language and setting state what the values were measured on, where given. Values may come from separate evaluation runs.

Key metric_comparison · version 2d581473-7212-4ea2-bf80-f0c4b8cb247e

Agentic task suite

100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

Key agentic_task_suite · version 0ea2c3bc-bc27-4395-9371-ce71dbf2c962

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04

Exact references

related

64285728-7520-4c5e-9d7a-9dbbe2c874f7

The original model's <think> tag on English retrieval prompts.