DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs
arXiv:2609.28570v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.AI
DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs