ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.16284v1 Announce Type: new Abstract: Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific eviden…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-16 04:00 · arXiv cs.CV
ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement