AINewsnow

Final-year student in India trying to break into generative-model inference optimization — roadmap feedback?

Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am: - Comfortable with C/C++ basics and PyTorch - Have done quantization work (GGUF/llama.cpp) - Working on a next-frame video prediction project (DiT + fl…

Read the full story at r/MLQuestions ↗

Timeline · 1 report

  1. 2026-09-30 06:54 · r/MLQuestions
    Final-year student in India trying to break into generative-model inference optimization — roadmap feedback?

More stories

  1. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  2. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  3. Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
  4. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  5. Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM
  6. Model Registry (RTX 4090) — r/LocalLLM
  7. dual 20gb 3080 and 4070ti local llm — r/LocalLLM
  8. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →