riderless: local decisions without text generation, using Gemma 4 26B-A4B on one 5090 (Apache-2.0)
I built riderless to get decisions out of a local model without generating an answer for the caller to parse. It runs a stock Gemma 4 26B-A4B on one RTX 5090 and returns choices, scores, and their distributions through an Apache-2.0 API, with zero generated tokens. The service is FastAPI over one l…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-22 06:05 · r/LocalLLM
riderless: local decisions without text generation, using Gemma 4 26B-A4B on one 5090 (Apache-2.0)
More stories
- Alibaba Unveils AI Chip to Drive Global Data Center Buildout — Bloomberg AI
- Amazon blocks Meta’s Muse AI agent — The Verge AI
- Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
- Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
- Lawsuit accuses Anthropic, OpenAI, SpaceXAI, Google of AI pacing 'collusion' — The Hill Technology
- Grok 4.7 — Hacker News Front Page
- Google's Gemini AI hacked three companies in security test — BBC Technology
- Meet the Data Agent in ChatGPT Work — OpenAI YouTube
Get the daily brief of stories like this at 6:30 every morning →