AINewsnow

Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the b…

Read the full story at r/computervision ↗

Timeline · 1 report

  1. 2026-10-07 09:18 · r/computervision
    Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

More stories

  1. Introducing Mistral Large 4 — Mistral AI News
  2. EmbeddingGemma 2: an open, lightweight multimodal embedding model — Google DeepMind Blog
  3. Sharing AI progress in mathematics — OpenAI News
  4. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  5. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  6. Boston Dynamics appoints Rohit Prasad as CEO; former Amazon AI chief to lead ‘Physical AI’ strategy — Mint AI
  7. ChatGPT for Teens is an ‘unacceptable risk,’ says Common Sense Media — The Verge AI
  8. Introducing the Decisions API — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →