H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder No SigLIP2 tower. No causal decoder. No VLM to repurpose. Here's how it works. π (1) Raw patches, not vision-tower features Images are split into non-overlapping 32Γ3β¦
Read the full story at r/machinelearningnews β
Timeline Β· 1 report
- 2026-09-06 21:15 Β· r/machinelearningnews
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder