PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO
arXiv:2610.04132v1 Announce Type: new Abstract: Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals,…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.CL
PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO