I have a rough idea about reducing repeated LLM inference across users — looking for technical feedback
This story is from 2026-08-21. It is preserved in the archive; the latest stories are on the live feed.
I'm fairly new to the deeper LLM inference/serving side, so I may be reinventing something that already exists. I'd really appreciate it if people here could point me toward existing work or explain where the idea breaks. The basic observation I had is: If 1,000 users ask different versions of esse…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-08-21 06:44 · r/deeplearning
I have a rough idea about reducing repeated LLM inference across users — looking for technical feedback