Prime announced the public release of Prime Inference, a serving platform that runs open‑source models on GPU clusters. The service supports both serverless and reserved capacity and includes automatic failover across data centers. Its first public endpoint, GLM‑5.3, was made available on OpenRouter on September 22 and has logged zero tool‑call errors and continuous uptime.
Prime Inference uses NVIDIA Blackwell GPUs, a stack built with Dynamo, vLLM, Mooncake and FlashInfer, and offers an OpenAI‑compatible API that can be called with any existing SDK.
Why it matters
Developers can now deploy open‑source models at production scale with reliable, low‑latency inference without building their own GPU infrastructure.