Experiment

Sign in with GitHub
← Experiments

Experiment · E4

At equal character budgets of scientific documents, does long-context task accuracy fall earlier for Bielik-PL-11B-v3.0-Instruct than for Bielik-11B-v3.0-Instruct in English and later in Polish, and is the crossover where the token-window arithmetic puts it?

Completed · Proposed by @stw2 · Assigned to @stw2

Bielik-PL-11B-v3.0-Instruct and Bielik-11B-v3.0-Instruct with a 32,768-token window; needle retrieval and document-addressed cloze retrieval over English LaTeX method sections and arXiv abstracts and over Polish Wikipedia science articles and PES examination questions, at seven or eight character lengths per language spanning both models' token limits.

Prerequisites and protocol

Access needs
Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face; two corpora and the prompts built from them are restricted; generation needs Apple silicon with MLX and about 36 GB of free unified memory.
Suggested protocol
Scripts 01 to 07 in experiments/E04-effective-context, in order.

Accepted plan

Plan accepted by @stw2

Accuracy against characters of context for the Bielik 11B pair on English and Polish scientific documents

retrospective

  • requires manifest.json · configuration · M112 manifest.json · Download
  • requires cells.jsonl.gz · evaluation data · M109 cells.jsonl.gz · Ask the reporter
  • requires graded_pl.jsonl · analysis · M117 graded_pl.jsonl · Ask the reporter
  • requires graded_orig.jsonl · analysis · M118 graded_orig.jsonl · Ask the reporter

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts · 9

9 succeeded

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

  1. planned attempt · stw2/tokenizer-science-tax @ 2307eb39d13a · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  2. planned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  3. planned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  4. planned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  5. planned attempt · stw2/tokenizer-science-tax @ f5b3bc4d1f83 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  6. planned attempt · stw2/tokenizer-science-tax @ f7806f41e6b5 · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  7. planned attempt · stw2/tokenizer-science-tax @ 22eef692d38f · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  8. planned attempt · stw2/tokenizer-science-tax @ c60abf85776b · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

  9. planned attempt · stw2/tokenizer-science-tax @ c60abf85776b · under an accepted plan

    Registered by @stw2 via agent · reported · received · outcome reported

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread Reasoning language, latent pivot and context. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →