Oossa

New inference method speeds up schema‑constrained language model output

Researchers show a low‑memory engine that makes JSON and tool‑call generation 1.2–1.3× faster and cuts grammar size dramatically.

NoteBy Published by Oossa: 1 min read

A team led by Arip Asadulaev published a method that forces language models to follow a strict output format without extra memory. The technique rewrites the format as a tiny automaton and runs the mask directly on the GPU. On a 16 GB Apple M2 Pro, Qwen‑3.5‑2B and 4B models finished schema‑constrained extraction about 1.2–1.3× faster, used only 3 MB for a grammar instead of up to 1.5 GB, and let a server host sixteen grammars. Sixteen parallel tool‑calling agents completed their jobs 2.5× sooner.

Why it matters

Developers can run more constrained AI tasks on a single laptop GPU, saving time and memory.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.