AINewsnow

GGUFs in transformers natively!

Hey there folks! Aritra here from Hugging Face. I wanted to update you all about the latest changes in `transformers`. We now natively support GGUFs (llama cpp quants). You can use it like so: from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "unsloth/Qwen3.5-4B-GGUF" filename…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-23 05:57 · r/LocalLLaMA
    GGUFs in transformers natively!

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. My contribution to the local AI community: 9 abliterated models, 99 GGUF quantizations in progress — r/huggingface
  3. Transformers now runs llama.cpp quants — Hugging Face Blog
  4. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM
  5. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  6. I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs — r/learnmachinelearning
  7. Don’t be fooled by this summer of AI hype — MIT Technology Review AI
  8. 2× Tesla P100 (2016 cards) in 2026: 110 tok/s on a 30B MoE, 16 tok/s at 1M context — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →