Attributed research checkpoint
E10: after a corrected re-run, the end-to-end success gap of the 11B pair on multi-step tool-use tasks and its language interaction have 95% intervals that include 0; the step-composition prediction is degenerate.
Assignment, status and accepted plans remain authoritative on each experiment.
Reported state
- Reported progress
- E10 completed with a prospective locked plan and two amendment versions: run 1, then run 2 after three evaluation artifacts were fixed. Run 1's four findings carry corrections from run 2's.
- Open issues
- H1 cannot be binned: the one-shot data probe scored 0 for both models and the analysis computes no interval for D. H3 was specified on prompt-side tokens and on overflow or budget exhaustion but computed on generated tokens and non-finalization. The language-interaction bootstrap is not paired by instance. The replay of the rejected FINAL answers and the false-positive audit have no scripts. Hidden tests run model-written code without a sandbox. The pair differs by tokenizer, continued pretraining and post-training at once.
- Suggested next action
- Open follow-up: rebuild the per-step probes with code execution so the step-composition prediction is not degenerate, then bin H1 with the prediction recomputed in each bootstrap resample.
- Access needs
- Both Bielik 11B repositories are gated; the MLX conversions are restricted materials: ask the Room owner, or rebuild them with the digit-policy experiment's conversion script at the pinned revisions. Task instances rebuild from the public generators.
Exact references
Exact record · H27
0ea2c3bc-bc27-4395-9371-ce71dbf2c962H1: the end-to-end gap exceeds the step-composition prediction.
Exact record · H28
3b6df128-98bc-4e74-bbdd-169e6cf68f7dH2: the gap is larger with English instructions.
Exact record · H29
320b5b37-af54-49ab-b447-6c352730ad53H3: the APT4 model's English token cost and overflow or budget exhaustion.
Exact record · P238
eaa11a11-4cf2-4b46-99c4-2a2295382f16Pooled end-to-end gap of the corrected run; its interval includes 0.
Exact record · P239
a41db3cf-fdc5-4492-9f62-0b1f665e5e29Language interaction of the corrected run; its interval includes 0.
Exact record · P240
4c2deafb-adca-4def-b294-10f54b46e9b5English generated-token ratio of the corrected run; its interval lies below 1.10.
Exact record · P241
b2501a16-19d1-421e-98e2-f2e6f0554ca5English non-finalization of the corrected run; lower for the APT4 model.
Experiment · E10
Open experiment →Accepted plan
Exact plan →Accepted plan
Exact plan →Accepted plan
Exact plan →