# Prime launches inference platform for open‑source models

> Prime Inference, a new serving platform for frontier open models, goes live with GLM‑5.3 on OpenRouter, offering low‑latency, 100% uptime.

Oossa · 2026-10-02 · https://oossa.com/en/prime-launches-inference-platform-for-open-source-models

Prime announced the public release of Prime Inference, a serving platform that runs open‑source models on GPU clusters. The service supports both serverless and reserved capacity and includes automatic failover across data centers. Its first public endpoint, GLM‑5.3, was made available on OpenRouter on September 22 and has logged zero tool‑call errors and continuous uptime.

Prime Inference uses NVIDIA Blackwell GPUs, a stack built with Dynamo, vLLM, Mooncake and FlashInfer, and offers an OpenAI‑compatible API that can be called with any existing SDK.

## The facts

- Prime Inference launched publicly on September 22, 2026 with the GLM‑5.3 model.
- The service runs on NVIDIA Blackwell GPUs and provides 100 % uptime since launch.

## Why it matters

Developers can now deploy open‑source models at production scale with reliable, low‑latency inference without building their own GPU infrastructure.

## Sources & references

1. [AnnouncementsOCT 02ND, 2026Prime Inference: Fast, Reliable Serving for Frontier Open Models](https://www.primeintellect.ai/blog/pi-inference-launch) – Prime Intellect
2. [AnnouncementsOCT 02ND, 2026Prime Inference: Fast, Reliable Serving for Frontier Open Models](https://www.primeintellect.ai/blog/prime-inference) – Prime Intellect

Last updated: 2026-10-02
