Help in setting up Pi-Agent
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
I set up a Qwen 4B model with Pi-Agent but the token output is decent but not instant (not expecting that but yea) I am using flash attention and the MTP with n gram spec set to 3 tokens. Any more suggestions to improve this setup would be highly appreciated !! Thankss !! :) EDIT: My bad for not pr…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-10 17:51 · r/LocalLLM
Help in setting up Pi-Agent