Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis
arXiv:2610.02718v1 Announce Type: new Abstract: Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-ce…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-05 04:00 · arXiv cs.CV
Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis