The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.15991v1 Announce Type: new Abstract: Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-16 04:00 · arXiv cs.CL
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization