OpenAI has added a new performance mode called Astra Ultrafast to its GPT‑6 model. It runs on NVIDIA’s Blackwell graphics chips and is already accessible through the OpenAI API for users of the ChatGPT Work and Codex plans. The company says the mode can generate tokens up to eight times faster than the standard Astra setting, making interactive AI applications feel more responsive.
Why speed matters for developers
Faster token generation shortens the loop when an AI‑assistant writes code, calls a tool, checks the result and decides the next step. OpenAI’s inference lead, Philippe Tillet, notes that the speed boost helps agents finish edit‑test‑debug cycles more quickly. The acceleration also reduces latency costs, which can matter when many calls are made in a single workflow.
How NVIDIA’s hardware supports the gain
OpenAI credits NVIDIA’s “deep investment in tooling and documentation” for allowing them to write high‑performance kernels that exploit the Blackwell architecture. The company also uses its own models to continuously improve the inference software that runs on the GPUs, meaning performance may keep improving after launch.
Why it matters
Developers who build code‑generation tools can see noticeably quicker responses, speeding up the edit‑test‑debug cycle. Faster loops also lower the time users wait for AI‑driven features in apps, improving the overall experience. How much the lower latency will reduce costs for end‑users remains to be seen.