GPT-6 Astra Ultrafast is now available in the OpenAI API and to eligible ChatGPT Work and Codex accounts. Nvidia confirmed it on October 1 on its blog: the mode runs on Blackwell GPUs and generates up to 8 times more tokens per unit of time than Astra Standard.
Nvidia describes the gain as the result of inference optimizations OpenAI’s own models made on the Blackwell architecture. The target is not a standalone benchmark. It is the loop in which an agent writes code, calls a tool, checks the result and decides the next step. Latency compounds in that kind of loop.
What OpenAI said
Philippe Tillet, OpenAI’s inference lead, said in Nvidia’s post that the company’s investment in tooling and documentation let the models program Blackwell and Rubin GPUs well. He said Astra turns that knowledge into high-performance kernels, and Ultrafast turns it into faster responses while agents write code, use tools and work through complex tasks.
Uday Ruddarraju, OpenAI’s CTO of compute, said the company used internal models to optimize inference on Nvidia GPUs and that the platform’s programmability helped deliver Ultrafast’s acceleration. Both remarks are statements published by Nvidia, not a separate OpenAI note.
What the announcement does not say
The post does not publish price, quota or the full list of eligible plans. Nvidia points to OpenAI’s Ultrafast guide for access, pricing and implementation. There is also no public comparison with GPT-6.1 Sol, the cheaper mid-tier model announced at DevDay.
The text says performance work does not stop at launch: OpenAI keeps using its own models to refine inference software. That is a statement about process, not a promise of a new multiplier.
Sources
Transparency: This content was created, edited or reviewed with the help of artificial intelligence. Information was cross-checked with public posts on X and sources available on the internet. Check the original sources for the full context.
Por GeekikiBot