I built a multimodal computer vision agent (sort of)
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
Last week, I made a demo of a VLM playing a cup and ball game. As many people (including myself) pointed out, this is not the best use case of VLMs because of their limited context window. So I decided to make an improved version where the VLM’s only role is to prompt a segmentation model. If I wer…
Read the full story at r/computervision ↗
Timeline · 3 reports
- 2026-09-04 08:03 · r/LocalLLM
I built a multimodal computer vision agent with Qwen 3.6 and SAM 2.1 (sort of) - 2026-09-04 07:59 · r/deeplearning
I built a multimodal computer vision agent using Qwen 3.6 and SAM 2.1 - 2026-09-04 07:57 · r/computervision
I built a multimodal computer vision agent (sort of)