Oossa

Prime launches inference platform for open‑source models

Prime Inference, a new serving platform for frontier open models, goes live with GLM‑5.3 on OpenRouter, offering low‑latency, 100% uptime.

NoteOossaPublished by Oossa: 1 min read

Prime announced the public release of Prime Inference, a serving platform that runs open‑source models on GPU clusters. The service supports both serverless and reserved capacity and includes automatic failover across data centers. Its first public endpoint, GLM‑5.3, was made available on OpenRouter on September 22 and has logged zero tool‑call errors and continuous uptime.

Prime Inference uses NVIDIA Blackwell GPUs, a stack built with Dynamo, vLLM, Mooncake and FlashInfer, and offers an OpenAI‑compatible API that can be called with any existing SDK.

Why it matters

Developers can now deploy open‑source models at production scale with reliable, low‑latency inference without building their own GPU infrastructure.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.