AINewsnow

Advice on models for RAG use case

Hi all, I'm looking for some advice picking models. I'm looking to try a project to ingest quite a large amount of docments into a knowledgebase, allowing me and others to ask questions about the data. I'm thinking of using Open WebUI and oikb for the interface and data ingestions, and llama.ccp/vl…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-27 09:11 · r/LocalLLaMA
    Advice on models for RAG use case

More stories

  1. Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding — MarkTechPost
  2. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  3. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
  4. JiRackUltra_1b Runs AI Routing on Any Laptop Without a GPU — AlphaSignal
  5. Targeting Applied AI / ML Engineer roles. Need ruthless feedback on my architecture and metrics. — r/deeplearning
  6. 42x Faster Prompt Lookup Drafting in llama.cpp — r/LocalLLaMA
  7. Best native alternative to WebUI for remote access to local LLMs? — r/LocalLLM
  8. Qwen3.8 27B on Intel X7 358h + B390, with pi + llama.cpp surprised by its own RAM speed — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →