Finding

Sign in with GitHub
← Publications

Finding · P204 · Author-curated

On 766 of 800 Polish STEM questions with Polish chain-of-thought, the logistic coefficient of answer correctness of Bielik-PL-11B-v3.0-Instruct on its standardised band English share, with benchmark fixed effects and standardised generated length, is 0.0907 log-odds per standard deviation (95% interval -0.1276 to 0.3386, bootstrap p-value 0.42, Holm-adjusted 0.42): no association is detected, and the interval does not exclude a positive association up to 0.3386.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Estimate with interval

metric
Pivot-accuracy coefficientQuantity estimated.
subject
Bielik-PL-11B-v3.0-InstructModel measured.
evaluation items
Polish STEM question setQuestion set.
items
766 questionsQuestions used.
value
0.0907 log-odds per standard deviationCoefficient estimate.
interval low
-0.1276 log-odds per standard deviationLower bound of the 95% bootstrap interval.
interval high
0.3386 log-odds per standard deviationUpper bound of the 95% bootstrap interval.
p value
0.42 p-valueTwo-sided bootstrap p-value.
setting
answers with Polish chain-of-thought; logistic regression by IRLS with benchmark fixed effects and standardised generated length; two-sided stratified paired item bootstrap, 2000 resamplesModel, covariates and test.
holm p
0.42 p-valuep_value after Holm correction over tests A1, A2 and B.

Experimental provenance

Method and evaluation protocol
Logistic regression by IRLS over the questions of answer correctness (from the graded outcomes) on the standardised band English share, benchmark indicators and the standardised number of re-encoded generated tokens.
Dataset
Greedy answers of both models to 800 Polish STEM questions (250 GSM8K-PL, 300 LLMzSzŁ STEM, 250 PES) with Polish chain-of-thought, and their graded outcomes; questions usable in both models after the lens exclusions.Version: unspecified · Access: restricted
Reported results
0.0907 [-0.1276, 0.3386], bootstrap p-value 0.42, Holm-adjusted 0.42.
Uncertainty and replication
95% stratified paired item bootstrap interval (2000 resamples, strata benchmark, seed 20260704), refitting the regression in each resample. Sensitivity: truncation-excluded 0.1297 [-0.1019, 0.3612] (741 items); compliance-restricted 0.0692 [-0.1628, 0.3078] (733 items); agreement at least 0.98 0.1328 [-0.1894, 0.4596] (498 items).
Limitations
The design registered no decision rule or equivalence bound for this test. The association is observational; nothing intervened on the pivot.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

Estimate with interval

The metric takes value on the evaluation_items, of which items are used, for the subject (relative to the comparator where one is given), with a 95% bootstrap interval from interval_low to interval_high and a two-sided p_value from the test named in setting; holm_p, where given, is that p-value after Holm correction; accuracy_english_cot and accuracy_polish_cot, where given, are the accuracies the metric is computed from.

Key estimate_with_interval · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Polish STEM question set

800 questions posed in Polish: 250 GSM8K-PL problems, 300 LLMzSzŁ STEM examination questions and 250 PES examination questions.

Key polish_stem_questions · version a8a88632-e5bd-42b6-8a04-17b77ce87d13

Pivot-accuracy coefficient

In a logistic regression over questions of a model's answer correctness on its standardised band English share, with benchmark fixed effects and the standardised number of generated tokens as covariates, the coefficient of the band English share in log-odds per standard deviation.

Key pivot_accuracy_coefficient · version a22bac3d-4e5b-4a59-b7c9-df5db8161a1e

Bielik-PL-11B-v3.0-Instruct

The 11B instruction-tuned Bielik v3 PL model with the APT4 tokenizer.

Key bielik_pl_11b_v3_instruct · version 55d90aa4-ac59-4189-bedf-b01cf526fe04