# Helion kernel adds to vLLM linear backend for faster inference

> vLLM now includes a Helion‑based linear backend that auto‑tunes matrix multiplication kernels, giving up to 10% higher throughput on Hopper GPUs.

Oossa · 2026-10-02 · https://oossa.com/en/helion-kernel-adds-to-vllm-linear-backend-for-faster-inference

The vLLM inference framework has integrated Helion, a PyTorch‑native kernel DSL, into its linear backend. Helion lets a single GEMM kernel cover multiple algorithm variants and automatically picks the fastest configuration for each shape. On NVIDIA Hopper GPUs the new backend beats vLLM’s default CUTLASS and DeepGEMM kernels, delivering consistent end‑to‑end speed gains and more than a 10% boost in throughput for some models.

## The facts

- Helion backend outperforms CUTLASS and DeepGEMM on Hopper GPUs, with >10% throughput improvement for certain workloads.
- The hybrid dispatch uses Helion for token counts up to 32 and falls back to default kernels for larger shapes.

## Why it matters

Developers running large language models on Hopper GPUs can see faster inference without writing custom kernels.

## Sources & references

1. [Building a High-Performance and Portable vLLM Linear Backend with Helion](https://pytorch.org/blog/building-a-high-performance-and-portable-vllm-linear-backend-with-helion/) – PyTorch

Last updated: 2026-10-02
