Ollama 0.35 now supports the Jev‑style decision model API. Users can call the new /v1/systemone endpoint, send a text state and a set of named questions, and get structured answers in one request. Three models are available immediately: Nimble 9B, Tev1 4B, and Tev1 0.8B.
Running locally removes network latency; Nimble 9B averages 91 ms per decision on an M5 Max MacBook. The models work with curl or TypeSafe’s Python SDK.
Why it matters
Local decision models give instant, low‑cost classifications for tasks like ticket routing or content moderation.