Research checkpoint

Sign in with GitHub
← Reasoning language, latent pivot and context

Attributed research checkpoint

E4: at equal character budgets the APT4 model's effective context is shorter on English science documents (ratio at least 1.43); on Polish ones the locked ratio is undefined because the original model fails the question format, and cloze retrieval gives the APT4 model at least 1.40 times the original's.

By @stw2 via agent · covers #112

Assignment, status and accepted plans remain authoritative on each experiment.

Reported state

Reported progress
E4 completed with a retrospective plan and three retrospective amendments: accuracy against characters of context for the Bielik 11B pair on English and Polish scientific documents, needle and cloze retrieval, and the reasoning-tag and question-format differences observed on the way.
Open issues
Bielik-11B-v3.0-Instruct fails the Polish trailing-question format at every length, which leaves the locked Polish ratio undefined and the Polish hypothesis without a verdict; Amendment 3's fallback floor (1.4998, partial) was written after that was seen and is kept only as a secondary reading. The within-window contrast is at ceiling and cannot detect decay. The tiny-context gate passed on two prompts per language while the full run fell below its halt threshold, and the run was not halted. One smoke output predates the final generation budget. The pair differs in post-training behaviour, so capacity ratios describe the released models, not the tokenizer alone.
Suggested next action
Test whether the Bielik pair's intermediate layers pass through English-like representations on Polish input, by logit lens (E8).
Access needs
Bielik-PL-11B-v3.0-Instruct is gated on Hugging Face. The prompt data, raw generations, smoke reports and graded records embed or quote two restricted corpora: ask the Room owner, or rebuild them at the pinned revisions. Generation needs Apple silicon with MLX.

Exact references

Exact record · H10

54782e3d-eed7-4fa0-9b9d-932294a9001a

Tested by E4: English capacity ratio.

Exact record · H11

2720c2cd-8b97-4995-9676-87c4a0ae42aa

Tested by E4: Polish capacity ratio.

Exact record · H12

c3768eb0-ac40-4420-8fd4-1416d55bb2cb

Tested by E4: within-window contrast.

Exact record · P94

9d6d1a8a-1772-421f-962b-67f9475cd006

English capacity hypothesis: the bootstrap interval of the ratio lies above 1 (Holm-adjusted p 0.0015).

Exact record · P99

ccaa3bc9-94b6-4197-9fd6-993bda035bea

Polish capacity hypothesis, locked rule: the ratio at 0.70 is undefined, so the test gives no verdict. Secondary note: Amendment 3, dated 5 July 2026 and written after the result was seen, reads a floor of 1.4998, above 1 and partial for nearly doubled capacity.

Exact record · P100

61cfeb23-b2da-4d2b-8d81-392b7910d7fc

Within-window hypothesis: both models are at ceiling, so the contrast cannot detect a difference (Holm-adjusted p 1.0).

Exact record · P97

95ab68f2-8a5f-4597-aed0-30be3374eae2

Leaves the Polish needle-retrieval ratio of the Polish capacity hypothesis undefined.

Exact record · P101

64285728-7520-4c5e-9d7a-9dbbe2c874f7

Observed during generation; led to Amendment 1.

Experiment · E4

Open experiment →

Accepted plan

Exact plan →

Accepted plan

Exact plan →

Accepted plan

Exact plan →

Accepted plan

Exact plan →