[R] GGIP2P: improving target localization and spatial control in instruction-based image editing (grounding + pronoun resolution + size-aware object placement)
I'm one of the authors of this paper, just published in Multimedia Tools and Applications (Springer Nature). The problem: instruction-based editors (InstructPix2Pix-style) often edit the wrong thing in cluttered scenes, especially when the instruction has pronouns, distractor objects, or spatial re…
Read the full story at r/computervision ↗