Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management
arXiv:2609.31648v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are studied mainly in robotics, where visual observations and language instructions are mapped to physical actions. This paper introduces Energy Vision-Language-Action (EVLA), a controlled multimodal benchmark for i…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv cs.LG
Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management