AINewsnow

local semantic file search for Linux (Rust, llama.cpp, EmbeddingGemma 2)

indexes PDFs, Office files, scans, screenshots, audio and video, and finds them by content. A query like "when did they decide to freeze hiring" opens the recording at that second. * EmbeddingGemma 2 embeds text, images, video frames and 30 s audio windows into one vector space, run in-process via…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-11 01:52 · r/LocalLLaMA
    local semantic file search for Linux (Rust, llama.cpp, EmbeddingGemma 2)

More stories

  1. Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM — r/LocalLLaMA
  2. Tested Mellum2.1-12B-A2.5B on PI Coding Agent - surprisingly usable, but not great at one-shot projects — r/LocalLLaMA
  3. Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s — r/LocalLLaMA
  4. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  5. Qwen3.8-Flash-Next (125B) at ~100 tok/s on an M5 Ultra Mac Studio with llama.cpp — r/LocalLLM
  6. Success with Qwen3.8 27B GSQ-RCO-IQ3_S on 16GB VRAM — r/LocalLLM
  7. Created laya : Now Introducing a new 800 Million Param physics-based typed decision model with 73k context and image support — r/LocalLLaMA
  8. Nemotron 3 Super (120B) at 43 tok/s on one RTX 4090, 2.5× faster than llama.cpp — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →