# New inference method speeds up schema‑constrained language model output

> Researchers show a low‑memory engine that makes JSON and tool‑call generation 1.2–1.3× faster and cuts grammar size dramatically.

Oossa · 2026-10-08 · https://oossa.com/en/new-inference-method-speeds-up-schema-constrained-language-model-output

A team led by Arip Asadulaev published a method that forces language models to follow a strict output format without extra memory. The technique rewrites the format as a tiny automaton and runs the mask directly on the GPU. On a 16 GB Apple M2 Pro, Qwen‑3.5‑2B and 4B models finished schema‑constrained extraction about 1.2–1.3× faster, used only 3 MB for a grammar instead of up to 1.5 GB, and let a server host sixteen grammars. Sixteen parallel tool‑calling agents completed their jobs 2.5× sooner.

## The facts

- Published: 2026-10-08 ("PUBLISHED: Thu Oct 08 2026")
- Speed‑up: 1.2–1.3× faster extraction; 2.5× faster tool agents; grammar size 3 MB vs 1.5 GB

## Why it matters

Developers can run more constrained AI tasks on a single laptop GPU, saving time and memory.

## Sources & references

1. [Breaking the Space Barrier and its Application to Language Model Inference](https://arxiv.org/abs/2610.09139) – arXiv, 2026-10-08

Last updated: 2026-10-08
