Experiment · E4
At equal character budgets of scientific documents, does long-context task accuracy fall earlier for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct in English and later in Polish, and is the crossover where the token-window arithmetic puts it?
Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct with a 32,768-token window; needle retrieval and document-addressed cloze retrieval over English LaTeX method sections and arXiv abstracts and over Polish Wikipedia science articles and PES examination questions, at seven or eight character lengths per language spanning both models' token limits.
Prerequisites and protocol
- Access needs
- Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face; two corpora and the prompts built from them are restricted; generation needs Apple silicon with MLX and about 36 GB of free unified memory.
- Suggested protocol
- Scripts 01 to 07 in experiments/E04-effective-context, in order.
Accepted plan
Plan accepted by @stw2
Accuracy against characters of context for the Bielik 11B pair on English and Polish scientific documents- requires manifest.json · configuration · M112 manifest.json · Download
- requires cells.jsonl.gz · evaluation data · M109 cells.jsonl.gz · Ask the reporter
- requires graded_pl.jsonl · analysis · M117 graded_pl.jsonl · Ask the reporter
- requires graded_orig.jsonl · analysis · M118 graded_orig.jsonl · Ask the reporter
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts · 9
9 succeeded
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
A25 Measures the character ceilings, emits the grids and assembles all prompts, cloze items and no-context controls.
Succeededplanned attempt · stw2/tokenizer-science-tax @ 2307eb39d13a · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan
A28 Greedy generation by the PL model for the no-context controls and every prompt within its token limit.
Succeededplanned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan
A29 Greedy generation by the original model for the no-context controls and every prompt within its token limit.
Succeededplanned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ f7806f41e6b5 · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ 22eef692d38f · under an accepted plan
planned attempt · stw2/tokenizer-science-tax @ c60abf85776b · under an accepted plan
A33 Writes flat headline values from the build manifest, the graded records, the analysis and the reasoning-tag counts.
Succeededplanned attempt · stw2/tokenizer-science-tax @ c60abf85776b · under an accepted plan
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →