I made LingBot-World 2.0 1.3B run at real-time 16 FPS on one RTX 5090. 2.5x faster than SGlang, 1.9x faster than Nvidia
Hi y'all, lots of cool world models have dropped recently, but most don't run on consumer GPUs in real time. LingBot-World 2.0 1.3B runs at 6 FPS on an RTX 5090 GPU. I made the same model run on 16 FPS on the same card. Then, I benchmarked all other inference engines and found 6-8 FPS is the best t…
Read the full story at r/StableDiffusion ↗
Timeline · 1 report
- 2026-09-18 20:51 · r/StableDiffusion
I made LingBot-World 2.0 1.3B run at real-time 16 FPS on one RTX 5090. 2.5x faster than SGlang, 1.9x faster than Nvidia
More stories
- NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
- Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0) — r/huggingface
- Huawei details AI accelerator roadmap, pulls in next-generation Ascend NPUs by several quarters — FP4 performance of the Ascend 960PR doubles expectations — Tom's Hardware
- Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. — r/LocalLLaMA
- Running Qwen3.8-Flash-Next ~85GB GGUF on 2× RTX 3060 12GB: ~12 tok/s, 131k ctx, CPU MoE, and a 26.5k agent prompt — r/LocalLLM
- Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. — r/LocalLLaMA
- FREE AI TRAINING CREDIT — r/learnmachinelearning
- No one is surprised that Nvidia's Jensen Huang thinks AI fears are overblown. — The Verge AI
Get the daily brief of stories like this at 6:30 every morning →