Experiment · E16
Is the bits-per-byte residual of the APT4 arm over the original-tokenizer arm on English and formal texts permanent, or does it keep closing when both continued-pretraining arms of Qwen2.5-1.5B continue to 1B tokens and are annealed?
Both continued-pretraining arms of Qwen2.5-1.5B resumed from their constant-learning-rate state at 500,170,752 Qwen tokens to 1B tokens on the same document sequence, which stays under one epoch, followed by the WSD decay anneal of both final checkpoints; bits per byte on the ten evaluation texts, the digit probe and the multiple-choice probes at every 100M-token milestone and after the anneal. Training and scoring.
Prerequisites and protocol
- Access needs
- An exact resume needs each arm's optimizer and data-sampler state and the packed training streams; none is among the archived materials, and the digests of the training text and packed streams were not recorded: ask the Room owner whether they survive, otherwise rebuild the streams with scripts 01 and 02 and record that the resume is not exact. The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Training needs one CUDA GPU with 16 GB, serially, about 1.5 days per arm.
- Suggested protocol
- Amendment 2 of the control-arm design, adopted on 8 July 2026 and never launched, as its own experiment. Commit a design with a decision rule before resuming. Training: scripts/04_train.py of experiments/E05-control-arm with --target-tokens 1e9 for the original-tokenizer arm and then the APT4 arm, each from its run directory, then --anneal on both final checkpoints before any headline value; the APT4 pack caps that arm near 1.0B tokens. Scoring: scripts/06_score_checkpoint.sh at each milestone; byte-anchored windows are added alongside the canonical windows once the per-position scoring rig exists, and checkpoints stay re-scorable, so the extension does not wait for it.
Accepted plan
No accepted plan. Available work need not have a complete protocol or source commit.
Responsibility and reported status
Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.
Room members can take or report on this experiment. Visit the Room to request membership.
Related findings
Execution attempts
Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.
No attempt registered. Work status and findings are independent of attempts.
Responsibility and plan history
Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”
Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.
Research guide →