Generalized Multimodal Foundation Model
arXiv:2609.22107v1 Announce Type: new Abstract: Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single tasks, making it difficult to quickly adapt t…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.LG
Generalized Multimodal Foundation Model