Experiment · E24
What is the between-run variation of each endpoint reported for the continued-pretraining arms of Qwen2.5-1.5B, and which reported differences are smaller than their minimum detectable effect?
Per arm, three continuations from one shared state at 450M tokens with different data orders for the final 50M tokens, a lower bound on full-run variance, scored on every endpoint of the control-arm experiment; a numerics term from the MPS and CPU parity band of the conformance battery; a per-endpoint uncertainty budget and minimum detectable effect, with the rule that no effect smaller than its band is claimed.
Prerequisites and protocol
- Access needs
- The control-arm design keeps optimizer state only in the latest training state, not in the 100M-token milestone checkpoints, and the archive holds neither the optimizer states nor the packed streams: ask the Room owner which training states survive; a shared state at 450M tokens needs a resumed run. The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Six short runs on CUDA GPUs with 16 GB or more.
- Suggested protocol
- Preconditions: a free training GPU after the 1B extension, and the conformance battery's parity band. Commit a design before training. A lower bound on variance can only remove effects, never validate one. Training: scripts/04_train.py of experiments/E05-control-arm with a data-order seed per continuation; scoring with scripts/06_score_checkpoint.sh.
Accepted plan
No accepted plan. Available work need not have a complete protocol or source commit.
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
No attempt registered. Work status and findings are independent of attempts.
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →