Experiment · E17
Do the scoring paths used in this Room satisfy model-independent equalities that any correct scorer must satisfy?
The bits-per-byte scorers, the likelihood multiple-choice scorers and the patching harness of the experiment folders, each run on a canary corpus of pathological text (no-break space, narrow no-break space, minus sign, tabs, newlines, indented code) against six relations: R1, decoding the encoding of a string returns the string; R2, per-window negative log-likelihood is invariant to batch sizes 1, 4 and 32 within 1e-6; R3, option log-probabilities are equivariant under option permutation; R4, results on MPS and on CPU agree within a declared, recorded band; R5, scoring from token ids equals scoring from strings; R6, a tokenizer that defines a beginning-of-sequence token places exactly one, at position 0. Tooling and gate runs on Apple silicon, 128 GB unified memory; no training.
Prerequisites and protocol
- Access needs
- The APT4 arm checkpoints used as positive controls are restricted materials: ask the Room owner. The APT4 tokenizer is gated on Hugging Face.
- Suggested protocol
- Commit a design listing each scoring script and the relations that apply to it. Each relation holds for a correct harness whatever the model, so a violation proves a defect. Positive controls the battery must flag: a bare AutoTokenizer load of an APT4 arm checkpoint directory, which deletes spaces and newlines at encoding (control-arm design, Amendment 1; embedding-initialisation erratum), and the MPS drift of the prefix-only scoring path recorded in the activation-patching report. Output: one signed gate file per scoring path; analysis scripts refuse to run without a passing gate file. The parity band of R4 is the numerics term of the variance-floor follow-up.
Accepted plan
No accepted plan. Available work need not have a complete protocol or source commit.
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
No attempt registered. Work status and findings are independent of attempts.
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →