AINewsnow

TileLang: Writing LLM GPU Kernels by Thinking in Tiles

This story is from 2026-09-22. It is preserved in the archive; the latest stories are on the live feed.

A modern LLM can spend most of its time doing something that looks almost embarrassingly simple: C = A @ B The mathematics is simple. Making an NVIDIA H100 execute that multiplication efficiently is a different problem. You have to decide which pieces of A and B live in HBM, which are copied into s…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-22 19:43 · DEV Community — AI
    TileLang: Writing LLM GPU Kernels by Thinking in Tiles

More stories

  1. Alibaba Unveils New AI Chip, Outlines Plan for Larger Model — Wall Street Journal Technology
  2. NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development — NVIDIA Blog
  3. NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories — NVIDIA Blog
  4. Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing — NVIDIA Technical Blog
  5. Topology-Aware Workload Scheduling with NVIDIA Topograph — NVIDIA Technical Blog
  6. Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS — NVIDIA Technical Blog
  7. Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton — NVIDIA Technical Blog
  8. 5 Companies Using NVIDIA AI for Clean Energy — NVIDIA Blog

Get the daily brief of stories like this at 6:30 every morning →