AINewsnow

ExLlamaSharp v1.4.0-beta: what shipped

This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.

ExLlamaSharp v1.4.0-beta is out. Local LLM server for Windows with NVIDIA GPUs — OpenAI-compatible /v1 , Blazor admin, and EXL3 inference. Release notes ## ExLlamaSharp 1.4.0-beta Multi-GPU (highlight) Real pipeline and tensor parallelism via ExLlamaV3 worker (Settings → Multi-GPU) Configurable Gpu…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-10 18:47 · DEV Community — AI
    ExLlamaSharp v1.4.0-beta: what shipped

More stories

  1. Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. — r/LocalLLaMA
  2. AI safety fears: King Charles meets tech leaders, here’s what he told OpenAI, Nvidia, Anthropic, DeepMind leaders — Mint AI
  3. Elon Musk talks up AI safety while fighting regulation in wild week of strange alliances — CNBC Technology
  4. King Charles warns of 'existential danger' of AI falling into wrong hands — BBC Technology
  5. Trump and Xi already talked AI guardrails once this year. Nothing was signed, and no chips shipped. — The Next Web
  6. I built a small proxy that lets Claude Desktop / Claude Code run on local models and NVIDIA's free API, sharing it in case it's useful — r/LocalLLM
  7. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  8. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology

Get the daily brief of stories like this at 6:30 every morning →