PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
arXiv:2609.36199v1 Announce Type: new Abstract: Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions.…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-30 04:00 · arXiv cs.CV
PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents