Finding

Sign in with GitHub
← Publications

Finding · P243 · Author-curated

Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct both succeed on none of the one-shot CSV data probes in English or Polish, and every task template has a step scored by that probe, so the step-composition prediction assigns both models a predicted success of 0 on every template, a predicted gap of 0 and no defined compounding factor.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Prediction floored by a probe

prediction
Step-composition predictionPrediction concerned.
subject
Bielik-PL-11B-v3.0-InstructModel after the tokenizer change.
comparator
Bielik-11B-v3.0-InstructModel before the tokenizer change.
benchmark
Agentic task suiteTasks predicted.
floored probe
one-shot data probe: a CSV-analysis question with the CSV inlined, answered without code executionPer-step probe with a success rate of 0.
predicted success
0 proportionPredicted success of both models on every template.
predicted gap
0 proportionComparator's minus subject's mean predicted success.
compounding factor defined
falseWhether the observed gap divided by the predicted gap is defined.

Experimental provenance

Method and evaluation protocol
Map each template's annotated steps to probe types (data-read and report to the data probe, code to the code probe, extract to the extract probe), take each model's one-shot success rate per probe and language, and compose the prediction.
Dataset
100 synthetic task instances from 20 parametric templates, each with English and Polish instructions.Version: tasks manifest sha256 66029e45c11f07613d698a67498b2c2c29faee4f0acc30bcf667bc3af6306f50 · Access: public
Reported results
Data probe success 0 for both models in both languages; extract probe Bielik-11B-v3.0-Instruct en 0.1667, Bielik-11B-v3.0-Instruct pl 0.1667, Bielik-PL-11B-v3.0-Instruct en 0.1667, Bielik-PL-11B-v3.0-Instruct pl 0.1; code probe Bielik-11B-v3.0-Instruct en 0.1667, Bielik-11B-v3.0-Instruct pl 0.1667, Bielik-PL-11B-v3.0-Instruct en 0.3333, Bielik-PL-11B-v3.0-Instruct pl 0.1667. Predicted gap 0; observed gap 0.025; D 0.025. The probes of run 1 are byte-identical.
Uncertainty and replication
No interval: the analysis computes no bootstrap for the prediction or for D.
Limitations
The locked design specified 30 code, 30 extract and 20 small data probe items plus reused arithmetic probes; the scripts pose every task instance one-shot instead, with CSVs of 60 to 400 rows. Single greedy run per cell on synthetic tasks; the two models differ by tokenizer, continued pretraining and post-training at once.

Author’s note

Values from results/static_probs.json and results/analysis.json (H1) at tag e10-run2-results; template steps from configs/tasks_manifest.json.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Prediction floored by a probe

Both models' success rate on floored_probe is 0 in both instruction languages and every template has a step scored by that probe, so the prediction gives predicted_success for both models on every template, predicted_gap is their difference and compounding_factor_defined states whether the observed gap divided by the predicted gap is defined.

Key prediction_floored · version c5c149a4-8643-4ec1-bf46-5e73abb0ef19

Agentic task suite

100 synthetic task instances, 5 from each of 20 parametric templates in three families (8 CSV analysis, 6 reproduction of a computation from a methods paragraph, 6 write-and-debug), each posed with English and with Polish instructions: 200 task cells per model.

Key agentic_task_suite · version 0ea2c3bc-bc27-4395-9371-ce71dbf2c962

Bielik-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 model with the original, Mistral-derived tokenizer.

Key bielik_11b_v3_instruct · version ac435442-b907-442a-9d1c-f951fa41d53c

Step-composition prediction

Predicted end-to-end success of a model on a template in one language: the product over the template's annotated steps of 1 - (1 - p)^(1 + r), where p is the model's one-shot success rate on the probe for the step's type in that language and r = floor((8 - number of steps) / number of steps). The predicted gap is the comparator's mean predicted success minus the subject's.

Key step_composition_prediction · version 0ea2c3bc-bc27-4395-9371-ce71dbf2c962

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04