AINewsnow

Model Registry (RTX 4090)

I like testing local models so I’m making a registry just for people who might be interested or care to discuss what works for them. You can see everything at Aether (your agent can browse it too) RTX 4090 Local Model Registry 24GB GPU: RTX 4090 - 24,564 MiB VRAM Runtime: llama.cpp / llama-swap KV:…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-30 03:43 · r/LocalLLM
    Model Registry (RTX 4090)

More stories

  1. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  2. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  3. Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
  4. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  5. Final-year student in India trying to break into generative-model inference optimization — roadmap feedback? — r/MLQuestions
  6. Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM
  7. dual 20gb 3080 and 4070ti local llm — r/LocalLLM
  8. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →