# APEX speeds up Qwen3-8B decoding up to 5.2×

> The new APEX controller cuts inference time for the Qwen3-8B model by up to 5.24 times, while lowering wasted draft tokens by 41%.

Oossa · 2026-10-07 · https://oossa.com/en/apex-speeds-up-qwen3-8b-decoding-up-to-5-2

Researchers introduced APEX, a learned controller that chooses how many draft tokens to generate and which speculation method to use. It plugs into the vLLM inference engine and was tested on the Qwen3-8B language model. On six benchmark workloads, APEX achieved up to 5.24× speedup over standard autoregressive decoding. The APEX‑Depth variant kept the speedup at 3.27× while cutting wasted tokens by 41% compared with a fixed n‑gram draft of size 16.

## The facts

- Up to 5.24× speedup on Qwen3-8B (APEX-R)
- 41.0% reduction in wasted tokens versus fixed n‑gram k=16 (APEX-B)

## Why it matters

Developers can run large language models faster and cheaper, making real‑time AI features more affordable.

## Sources & references

1. [APEX: Speculate smarter, not deeper](https://arxiv.org/abs/2610.07780) – arXiv, 2026-10-07

Last updated: 2026-10-07
