AINewsnow

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp. We're both software engineers and previously built an o…

Read the full story at Hacker News Front Page ↗

Timeline · 1 report

  1. 2026-09-30 17:37 · Hacker News Front Page
    Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

More stories

  1. If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context. — r/LocalLLaMA
  2. What model sits between Qwen 3.8 27b and Flash next for coding? — r/LocalLLaMA
  3. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Why the same Llama 3.2 1B model comes in different file sizes: a beginner’s explanation — r/AI_Agents
  5. Qwen3.8-Flash-Next on 12GB VRAM - 65 t/s — r/LocalLLM
  6. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  7. TIL about llama.cpp's RPC (Remote procedure call), might be better than Vulkan? YMMV — r/LocalLLM
  8. Strata Qwen3.8 FN abliterated dual 16GB GPU + 128GB RAM — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →