Ornith-1.5-35B-A3B (Q4, 8-bit context) ~44 tok/s on an RTX 3050 8GB
Lamina runs 35B-A3B MoE models (Qwen3.6, Ornith) locally on Windows with 8 GB VRAM by keeping hot experts on the GPU and streaming the rest from system RAM. 128K context, MTP speculative decoding, vision input, OpenAI-compatible server. ~44 tok/s on an RTX 3050 8GB (Ryzen 7 5700X) vs ~28.6 tok/s fo…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-09 21:14 · r/LocalLLM
Ornith-1.5-35B-A3B (Q4, 8-bit context) ~44 tok/s on an RTX 3050 8GB