Would you use a remote H3 text-encoder API so the 32B stays off your GPU?
H3 local users already know the pain: the DiT wants ~20–25GB, the encoder is a truncated Qwen3-VL-32B, and they do not co-fit on a 24–32GB card. Leave the encoder loaded and the sampler streams. Unload it and you wait on the next prompt. I’m considering a text-only encode API for T2VA: You send the…
Read the full story at r/comfyui ↗
Timeline · 1 report
- 2026-09-28 18:10 · r/comfyui
Would you use a remote H3 text-encoder API so the 32B stays off your GPU?