Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.12454v1 Announce Type: new Abstract: Vision-Language Models such as CLIP enable effective few-shot medical anomaly detection (AD) via strong image-text semantic alignment. However, their globally contrastive pretraining lacks explicit spatial supervision, limiting precise lesion localiza…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-14 04:00 · arXiv cs.CV
Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images