Execution attempt · A68 · rerun
Re-scores the base model and the three unchanged transplants with whitespace-gated tokenizer loads after the erratum.
Pinned source and configuration
https://github.com/stw2/tokenizer-science-tax @ 6c8a461e502d64cdecbb61d981a99dac9addbb67
Reference checked 2026-09-14 13:12 UTC. Later commits, branches or plan changes do not retarget this attempt.
- Command
- python scripts/03_score_bpb.py --model base=Qwen/Qwen2.5-1.5B fvt=models/fvt focus=models/focus random=models/random
- Working directory
- experiments/E07-embedding-init
- Configuration paths
- None
- Parameters
- max-docs 200; seq 2048 tokens; add_special_tokens false; bf16; row-wise scoring; tokenizers through load_tokenizer(): whitespace canary, pristine APT4 fallback; rows into a fresh file; finish time not recorded, reported at the results commit
- Environment
- 3.12.13; PyTorch MPS, bf16, row-wise scoring; Apple silicon, 128 GB unified memory
- Output directory
- Not recorded
Inputs
Materials the registrant named when registering this attempt, by content identity; obtainability is derived from their location reports. Nothing is fetched or verified.
- Qwen2.5-1.5B · base model · M124 Qwen2.5-1.5B · checkpoint · Download
- models/fvt · model · M357 models/fvt · raw output · Ask the reporter
- models/focus · model · M358 models/focus · raw output · Ask the reporter
- models/random · model · M356 models/random · raw output · Ask the reporter
- APT4 tokenizer · reference tokenizer · M3 APT4 tokenizer (Bielik-PL-11B-v3.0-Instruct) · tokenizer · Ask the reporter
- pl: pl_eval.jsonl · evaluation data · M129 pl_eval.jsonl · raw output · Download
- en: en_eval.jsonl · evaluation data · M130 en_eval.jsonl · raw output · Ask the reporter
- pl_informal: pl-informal.jsonl · evaluation data · M33 pl-informal.jsonl · raw output · Ask the reporter
- pl_wiki_sci: pl-wiki-science.jsonl · evaluation data · M32 pl-wiki-science.jsonl · raw output · Download
- pl_pes: pl-science-pes.jsonl · evaluation data · M31 pl-science-pes.jsonl · raw output · Ask the reporter
- sci_arxiv: en-arxiv-abstracts.jsonl · evaluation data · M27 en-arxiv-abstracts.jsonl · raw output · Download
- sci_latex: en-latex-methods.jsonl · evaluation data · M28 en-latex-methods.jsonl · raw output · Ask the reporter
- sci_python: en-python-code.jsonl · evaluation data · M29 en-python-code.jsonl · raw output · Ask the reporter
- sci_gsm8k: en-gsm8k.jsonl · evaluation data · M30 en-gsm8k.jsonl · raw output · Download
- math_clean: math_clean.jsonl · evaluation data · M128 math_clean evaluation statements · dataset · Download
Delivered events
Registered
#1Re-scores the base model and the three unchanged transplants with whitespace-gated tokenizer loads after the erratum.
Started
#2Started.
Succeeded
#3Completed with exit code 0.
Exit code 0.
- bpb_e7.jsonl · https://github.com/stw2/tokenizer-science-tax/blob/e7d4bb2df26254f88ecbf390a7cbf69ac210e0e3/experiments/E07-embedding-init/results/bpb_e7.jsonl · public · sha256 cf4ce45e06d2… · 4780 bytes · M367
- perdoc/ · https://github.com/stw2/tokenizer-science-tax/tree/e7d4bb2df26254f88ecbf390a7cbf69ac210e0e3/experiments/E07-embedding-init/results/perdoc · public · sha256 1a554e8b9c74… · 148080 bytes · M368
Report an event
The registrant’s capture tool normally delivers start and outcome events. Reporting here is the same author report with the browser as the reported time; it does not observe the process.
This attempt has a delivered outcome. A new execution is a new attempt.