Cited claim

Sign in with GitHub
← Publications

Cited claim · P21 · Author-curated

FOCUS represents each token of the target vocabulary as a sparse linear combination of tokens from the original vocabulary, selected by semantic similarity in an auxiliary embedding space.

Published by @stw2 · 2026-09-14 · Sources, measurements and interpretation are supplied by the author.

Structured assertion

Relation: Method mechanism

method
FOCUSThe method.
mechanism
sparse linear combination of original-vocabulary tokens for each target-vocabulary tokenHow the method builds its output.
selection
semantic similarity in an auxiliary embedding spaceHow contributing tokens are chosen.

Paper citations

arXiv:2604.10799v1 →Revision supplied by author
  1. Section 4, p. 3

    The FOCUS method represents each token in the target vocabulary as a sparse linear combination of tokens from the original vocabulary, selected based on semantic similarity in an auxiliary embedding space.

Concept definitions

Reuse the defining version and key when the meaning fits your assertion.

FOCUS

Fast Overlapping Token Combinations Using Sparsemax: an embedding initialisation for a replaced tokenizer's vocabulary.

Key focus_init · version 5a078e71-84ff-4039-9e32-7999a5f679f5

Method mechanism

The method builds its output by the mechanism; selection states how inputs are chosen and consequence an effect on training, where given.

Key method_mechanism · version 5a078e71-84ff-4039-9e32-7999a5f679f5