AINewsnow

Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)

This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.

VLLM Benchmark: Prefill, Prompt processing - avg, 871.93 tok/s (3 hours constant running xhigh) - 10K prompt, 1000.26 tok/s (16 runs) - 90K prompt, 743,59 tok/s (16 runs) Decode, tok gen - avg, 38.39 tok/s (3 hours constant running xhigh) - 10K, 42.3 tok/s (16 runs) - 90K, 34 tok/s (16 runs) Preamb…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-13 17:25 · r/LocalLLaMA
    Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT)

More stories

  1. Trump announces a new 'AI Force,' but says he will not 'stifle' AI — Business Insider AI
  2. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  4. Introducing Astra for Law — OpenAI News
  5. Microsoft exec called AI scraping the “largest theft of labor in human history” — Ars Technica AI
  6. Newsom signs executive order to explore new AI rules, consider ‘kill switch’ — Politico Technology
  7. Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks — Bloomberg AI
  8. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology

Get the daily brief of stories like this at 6:30 every morning →