FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face
from FreedomIntelligence: HuatuoGPT-3-27B is a medical LLM built on Qwen3.8-27B with One-stage Policy Optimization (OnePO) . OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher responses provide tempora…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-24 19:20 · r/LocalLLaMA
FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face