tencent/WeMM-Embedding 9B/4B/2B
This story is from 2026-08-25. It is preserved in the archive; the latest stories are on the live feed.
WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported. https://huggingface.co/tencent/WeMM-Embedding-9B…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-25 09:36 · r/LocalLLaMA
tencent/WeMM-Embedding 9B/4B/2B