Java vllm-like framework claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile
TornadoVM: The Java to CUDA engine: https://github.com/beehive-lab/TornadoVM jitLLM: The inference engine: https://github.com/beehive-lab/jitllm Deep dive talk: https://www.youtube.com/watch?v=HO5CpETzywk
Read the full story at r/LocalLLaMA ↗