AINewsnow

Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

https://huggingface.co/Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit Qwen3.8-27B CODER — IQ4_XS imatrix · 24 GB card fit · ~262k context · MTP draft Quantized, Abliterated, and fitted by LexiPanel. Its Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k tokens of context.…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-02 13:32 · r/LocalLLaMA
    Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

More stories

  1. We just open-sourced the world's fastest WebGPU kernels for local AI on Hugging Face — r/LocalLLaMA
  2. Rogue AI agents: A timeline of security breaches since the attack on Hugging Face — Fast Company AI
  3. Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo — r/LocalLLM
  4. Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence — MarkTechPost
  5. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  6. Anyone tried Swift 1.5 Flash Next GSQ-RCO IQ2_XS + Strata? — r/LocalLLM
  7. Bytedance release 4-step for Minimax-h3; DMAD: Distribution Matching as Adversarial Distillation — r/StableDiffusion
  8. Nova v3: 148M model, 1% of the data, same benchmark scores as SmolLM2-135M — r/huggingface

Get the daily brief of stories like this at 6:30 every morning →