MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting
arXiv:2609.28857v1 Announce Type: new Abstract: Scene text spotting remains challenging for arbitrarily shaped text instances such as curved signs and dense multi-oriented characters in natural images, where tightly coupled architectures propagate localization errors directly into recognition failu…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.CV
MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting