A team led by Arip Asadulaev published a method that forces language models to follow a strict output format without extra memory. The technique rewrites the format as a tiny automaton and runs the mask directly on the GPU. On a 16 GB Apple M2 Pro, Qwen‑3.5‑2B and 4B models finished schema‑constrained extraction about 1.2–1.3× faster, used only 3 MB for a grammar instead of up to 1.5 GB, and let a server host sixteen grammars. Sixteen parallel tool‑calling agents completed their jobs 2.5× sooner.
Why it matters
Developers can run more constrained AI tasks on a single laptop GPU, saving time and memory.