Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU
This story is from 2026-10-07. It is preserved in the archive; the latest stories are on the live feed.
EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the b…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-07 09:26 · r/LocalLLaMA
Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU