Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
arXiv:2609.30288v1 Announce Type: new Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternati…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv cs.CL
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling