Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
arXiv:2610.10991v1 Announce Type: new Abstract: Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt down…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.CV
Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval