Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine
Maximizing perf on AI-SQL queries with the KV-optimal left-deep join
Read the full story at Modal Blog ↗
Timeline · 1 report
- 2026-09-24 00:00 · Modal Blog
Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine