Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models
arXiv:2609.28865v1 Announce Type: new Abstract: Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normal…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.CV
Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models