Model: add Tencent Hy 4 (hy_v4) preview architecture support by Little0o0 · Pull Request #28127 · ggml-org/llama.cpp
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
Model : https://huggingface.co/tencent/Hy4-preview
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-09-06 16:36 · r/LocalLLaMA
[Model] Support for Spark2_5ForCausalLM implementation by KnightYao · Pull Request #27868 · ggml-org/llama.cpp - 2026-09-06 02:16 · r/LocalLLM
KV Cache Streaming from RAM - 2026-09-06 02:15 · r/LocalLLaMA
Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant - 2026-09-04 13:20 · r/LocalLLaMA
Model: add Tencent Hy 4 (hy_v4) preview architecture support by Little0o0 · Pull Request #28127 · ggml-org/llama.cpp