AINewsnow

Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel

I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length. The results and code are below: Results (same hardware, my megakernel vs llama.cpp with MTP): - Writing code: 140 tok/…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-10 13:01 · r/LocalLLaMA
    Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel

More stories

  1. Found a fix for AMD RX 6600 100% CPU usage — r/LocalLLM
  2. Is it just my migraine or did every closed Ai get downgraded to Artificial Ignorance? — r/ArtificialInteligence
  3. Claude too expensive for companies, only 2.2% of US households paying for AI — r/ArtificialInteligence
  4. Jev for beginners: how to use it and what to build — How I AI
  5. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  6. Introducing GPT-6 in ChatGPT with Intelligent UI — OpenAI YouTube
  7. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  8. Anthropic bans users from ‘needless abusive or cruel behavior’ towards Claude — The Guardian AI

Get the daily brief of stories like this at 6:30 every morning →