Oossa

Helion kernel adds to vLLM linear backend for faster inference

vLLM now includes a Helion‑based linear backend that auto‑tunes matrix multiplication kernels, giving up to 10% higher throughput on Hopper GPUs.

NoteOossaPublished by Oossa: 1 min read

The vLLM inference framework has integrated Helion, a PyTorch‑native kernel DSL, into its linear backend. Helion lets a single GEMM kernel cover multiple algorithm variants and automatically picks the fastest configuration for each shape. On NVIDIA Hopper GPUs the new backend beats vLLM’s default CUTLASS and DeepGEMM kernels, delivering consistent end‑to‑end speed gains and more than a 10% boost in throughput for some models.

Why it matters

Developers running large language models on Hopper GPUs can see faster inference without writing custom kernels.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.