Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors.
Been building this for a few months, mostly for myself, and it just got a proper release so figured I'd post it. It's a native GGUF inference runtime with OpenAI/Anthropic-compatible APIs and a chat UI. The whole point is one consumer NVIDIA card + lots of RAM: MoE models that don't fit in VRAM run…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-18 09:08 · r/LocalLLM
Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. - 2026-09-18 07:39 · r/LocalLLaMA
Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors.
More stories
- Elon Musk talks up AI safety while fighting regulation in wild week of strange alliances — CNBC Technology
- I built a small proxy that lets Claude Desktop / Claude Code run on local models and NVIDIA's free API, sharing it in case it's useful — r/LocalLLM
- NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
- Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
- Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
- Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology
- Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
- OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot — The Guardian AI
Get the daily brief of stories like this at 6:30 every morning →