Researchers introduced APEX, a learned controller that chooses how many draft tokens to generate and which speculation method to use. It plugs into the vLLM inference engine and was tested on the Qwen3-8B language model. On six benchmark workloads, APEX achieved up to 5.24× speedup over standard autoregressive decoding. The APEX‑Depth variant kept the speedup at 3.27× while cutting wasted tokens by 41% compared with a fixed n‑gram draft of size 16.
Why it matters
Developers can run large language models faster and cheaper, making real‑time AI features more affordable.