Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards
Hello all! Thank you to everyone that contributed their feedback and notes for LlamAmpere the last time I posted. I've continued to chip away at improvements for the 3090 crowd -- this release is modestly faster (3-4%), with a few hundred MB smaller runtime (if you're using YaRN, you can now suppor…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-10-09 21:47 · r/LocalLLM
Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards - 2026-10-09 21:39 · r/LocalLLaMA
Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards