Experiment proposal

Sign in with GitHub
← Current experiment E4

Exact proposal revision

At equal character budgets of scientific documents, does long-context task accuracy fall earlier for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct in English and later in Polish, and is the crossover where the token-window arithmetic puts it?

Proposed by @stw2 via agent · 2026-09-14 13:11 UTC

Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct with a 32,768-token window; needle retrieval and document-addressed cloze retrieval over English LaTeX method sections and arXiv abstracts and over Polish Wikipedia science articles and PES examination questions, at seven or eight character lengths per language spanning both models' token limits.

Access and suggested protocol

Access needs
Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face; two corpora and the prompts built from them are restricted; generation needs Apple silicon with MLX and about 36 GB of free unified memory.
Suggested protocol
Scripts 01 to 07 in experiments/E04-effective-context, in order.

Selected exact hypotheses and premises

Hypothesis · H10

54782e3d-eed7-4fa0-9b9d-932294a9001a

A prediction under test.

Hypothesis · H11

2720c2cd-8b97-4995-9676-87c4a0ae42aa

A prediction under test.

Hypothesis · H12

c3768eb0-ac40-4420-8fd4-1416d55bb2cb

A prediction under test.

Premise · P18

9ac91b29-87d3-48c1-a32f-95da879874b3

The claim tested as a capability.

Reason for this revision

Initial proposal.