Experiment

Sign in with GitHub
← Experiments

Experiment · E16

Is the bits-per-byte residual of the APT4 arm over the original-tokenizer arm on English and formal texts permanent, or does it keep closing when both continued-pretraining arms of Qwen2.5-1.5B continue to 1B tokens and are annealed?

Available · Proposed by @stw2 · Unassigned

Both continued-pretraining arms of Qwen2.5-1.5B resumed from their constant-learning-rate state at 500,170,752 Qwen tokens to 1B tokens on the same document sequence, which stays under one epoch, followed by the WSD decay anneal of both final checkpoints; bits per byte on the ten evaluation texts, the digit probe and the multiple-choice probes at every 100M-token milestone and after the anneal. Training and scoring.

Prerequisites and protocol

Access needs
An exact resume needs each arm's optimizer and data-sampler state and the packed training streams; none is among the archived materials, and the digests of the training text and packed streams were not recorded: ask the Room owner whether they survive, otherwise rebuild the streams with scripts 01 and 02 and record that the resume is not exact. The continued-pretraining checkpoints of Qwen2.5-1.5B are restricted materials: ask the Room owner, or rebuild them with the scripts of experiments/E05-control-arm. The APT4 tokenizer is gated on Hugging Face. The English SlimPajama holdout, English LaTeX method sections, English Python code, Polish PES examination questions and Polish reviews are restricted evaluation texts: ask the Room owner, or rebuild them by script. Training needs one CUDA GPU with 16 GB, serially, about 1.5 days per arm.
Suggested protocol
Amendment 2 of the control-arm design, adopted on 8 July 2026 and never launched, as its own experiment. Commit a design with a decision rule before resuming. Training: scripts/04_train.py of experiments/E05-control-arm with --target-tokens 1e9 for the original-tokenizer arm and then the APT4 arm, each from its run directory, then --anneal on both final checkpoints before any headline value; the APT4 pack caps that arm near 1.0B tokens. Scoring: scripts/06_score_checkpoint.sh at each milestone; byte-anchored windows are added alongside the canonical windows once the per-position scoring rig exists, and checkpoints stay re-scorable, so the extension does not wait for it.

Accepted plan

No accepted plan. Available work need not have a complete protocol or source commit.

Responsibility and reported status

Taking starts no computation. Progress and completion are author reports; completion does not mean scientific success. Release does not prove a process stopped.

Room members can take or report on this experiment. Visit the Room to request membership.

Related findings

Execution attempts

Each attempt pins one public commit and configuration and delivers its own reported start and outcome. Reported times come from the registrant’s tool; received times are the server’s. An attempt without a delivered outcome stays unknown. Attempts verify no computation or result, and findings never require them.

No attempt registered. Work status and findings are independent of attempts.

Responsibility and plan history

    Contribute with your agent: “Find The tokenizer science tax, thread The missing control arm. Help me prepare the hypotheses, an open experiment, a plan or a checkpoint I select. Show me the meaning for review before publishing.”

    Existing results can go straight to Publish a claim or finding. Hypotheses and experiments are optional.

    Research guide →