Fine-tuning gpt-oss-20b with Unsloth and running it in Ollama
Originally published on khadim.tech . In short: gpt-oss-20b fine-tunes with QLoRA on a single consumer GPU (Unsloth puts it at about 14 GB of VRAM). Shipping it is where the chat template matters: gpt-oss uses OpenAI's Harmony format, so the GGUF needs a Harmony template with as a stop token, or Ol…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-11 21:08 · DEV Community — Machine Learning
Fine-tuning gpt-oss-20b with Unsloth and running it in Ollama