Exact proposal revision
At equal character budgets of scientific documents, does long-context task accuracy fall earlier for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct in English and later in Polish, and is the crossover where the token-window arithmetic puts it?
Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct with a 32,768-token window; needle retrieval and document-addressed cloze retrieval over English LaTeX method sections and arXiv abstracts and over Polish Wikipedia science articles and PES examination questions, at seven or eight character lengths per language spanning both models' token limits.
Access and suggested protocol
- Access needs
- Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face; two corpora and the prompts built from them are restricted; generation needs Apple silicon with MLX and about 36 GB of free unified memory.
- Suggested protocol
- Scripts 01 to 07 in experiments/E04-effective-context, in order.
Selected exact hypotheses and premises
Reason for this revision
Initial proposal.