
1 Oct 2026
NVIDIA says GPT-6 Astra Ultrafast on Blackwell is up to 8× faster than Standard
On Oct. 1, 2026, NVIDIA said GPT-6 Astra Ultrafast — running on Blackwell GPUs with OpenAI-tuned inference software — is available now and can generate tokens up to eight times faster than Astra Standard.
Agent loops die on waiting. Shipping a named Ultrafast Astra mode on Blackwell — with OpenAI saying it used its own models to tune the kernels — is a concrete latency product, not another vague “AI factory” slogan.
On Thursday, 1 October 2026, NVIDIA said GPT-6 Astra Ultrafast is available now. The company blog is titled “How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast.” The byline is Dion Harris. The page dates the post October 1, 2026, and does not print an hour. The line under the title says that, running on NVIDIA Blackwell GPUs and accelerated by continuous inference optimizations through OpenAI’s models, Ultrafast delivers faster model responses across code generation, tool use, and interactive applications. The post says Ultrafast, running on those Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. A GPU is the chip that does the model’s math. Blackwell is the NVIDIA chip generation the post names. Inference is the step where a trained model writes an answer, as distinct from the earlier step of training it. An API is the connection a program uses to ask for that answer. Codex, here, is OpenAI’s coding product. Eligible is NVIDIA’s word for who can turn the mode on in ChatGPT Work and in Codex. The post does not list those eligibility rules. Those lines are NVIDIA’s.
What NVIDIA says the speedup is. Accelerated by inference optimizations through OpenAI’s models that use the Blackwell architecture, Ultrafast offers up to 8x faster token generation than Astra Standard mode. A token is a small piece of the text the model reads or writes. Token generation is the model writing those pieces. Up to eight times is the ceiling NVIDIA states. It is not a line that every request finishes eight times sooner. Astra Standard is the comparison mode the post names. For developers, NVIDIA says faster generation can shorten a coding agent’s edit-test-debug cycle, cut the time spent writing a reply between tool calls, and make an interactive application feel more responsive. An agent, here, is software that takes a next step, not only a chat reply. A tool call is the agent asking another program to do something, such as run a test. Those lines are NVIDIA’s.
Why NVIDIA says the wait matters when it repeats. A faster response matters most when it is repeated across a workflow: an agent writes code, uses a tool, checks the result, and decides what to do next. Ultrafast brings Astra’s capabilities into these time-sensitive loops. NVIDIA says its AI infrastructure helps OpenAI serve more useful model outputs when developers need them. Those lines are NVIDIA’s.
What OpenAI’s inference lead said, as NVIDIA quotes him. Philippe Tillet, inference lead at OpenAI, said NVIDIA’s investment in tooling and documentation has made OpenAI’s models exceptionally good at programming Blackwell and Rubin GPUs. He said Astra can turn that knowledge into high-performance kernels that make NVIDIA hardware compelling across latency, throughput, and cost. With Astra Ultrafast, he said, that means faster model responses as agents write code, use tools, and work through complex tasks. A kernel, here, is the small program that runs the model’s math on the chip. Latency is the wait. Throughput is how much work the machines finish in a stretch of time. Rubin is the GPU name he places next to Blackwell. The post does not give a date for Rubin hardware. That quotation is his, on the NVIDIA page.
What NVIDIA says happens after a model is already serving people. The section is titled “Continually Improving Performance.” Performance gains do not stop when a model is deployed. OpenAI is using its own models to help refine the inference software running on NVIDIA GPUs, and it is using the platform’s programmability to test and implement improvements. Deployed means the model is already answering real requests. Programmability, in that sentence, means the software on the chip can be changed, rather than only used as it first shipped. NVIDIA says that ongoing work can make model responses faster, and can make the machines already installed more productive over time. Those lines are NVIDIA’s.
What OpenAI’s compute chief said, as NVIDIA quotes him. Uday Ruddarraju, chief technology officer of compute at OpenAI, said the work with NVIDIA is helping OpenAI make AI faster and more useful. He said OpenAI used its internal models to optimize inference on NVIDIA GPUs, and that NVIDIA’s programmability helped deliver the acceleration behind Astra Ultrafast. Optimize, here, means change the software so the same chips answer faster. Internal models are models OpenAI runs itself. The sentence says those models were used to do the optimization. That quotation is his, on the NVIDIA page.
What NVIDIA says a programmable platform is for, past this one mode. The post says a programmable NVIDIA platform lets developers and researchers reuse the same machines across training, inference, and reinforcement learning as models change. Training is teaching the model. Reinforcement learning, in that sentence, is a loop that changes the model based on how its answers score. NVIDIA says the flexibility helps teams move compute to where the demand is, raise utilization, and avoid overprovisioning each kind of work. Compute is the machines. Utilization is how much of the reserved hardware is doing useful work. Overprovisioning means keeping more machines than that job needs. Those lines are NVIDIA’s. The post does not print a utilization percentage.
Where NVIDIA sends a developer for access, price, and how to turn it on. The post says developers can use GPT-6 Astra Ultrafast through the API today, and it points to an Ultrafast guide for access, pricing, and implementation details. The guide, on OpenAI’s site, calls Ultrafast the fastest service tier in the API, with up to 8x faster speeds than Standard mode, and says to use it when speed justifies the higher cost. Higher cost is the guide’s phrase. The NVIDIA post does not print a dollar price, and it does not say the mode is free. The guide says to see its pricing table for input, cached input, cache write, and output prices. The guide also says a persistent connection, which it calls a WebSocket, matters for an agent that makes many tool calls in a row, because the wait of opening a new connection each time can eat the latency gain. It says to set the model to gpt-6-astra and the service tier to ultrafast. It says Ultrafast for GPT-6 Astra is available to API users at low rate limits, and that an organization with an OpenAI account team can ask that team for a higher limit. A rate limit is a cap on how much text the account may send and receive, not the eight-times generation claim. The guide’s default caps, in tokens per minute, are 500,000 for usage tiers 1 through 3, 1 million for tier 4, and 5 million for tier 5. A token per minute is that cap. It is not a measure of how fast one reply is written. The guide says Ultrafast supports US data residency and global processing only, and does not support EU or other non-US regional processing endpoints. Data residency, here, means which region is allowed to handle the request. Those guide lines are OpenAI’s. The guide does not restate NVIDIA’s sentence about ChatGPT Work and Codex.
The picture is NVIDIA’s comparison graphic. The left panel is labeled Standard, with 17.0s under the word, and shows a white egg-shaped machine with a black window and two black arms on a round platform. The right panel has a green border, is labeled Ultrafast, with 6.4s under the word, and shows a white rocket with colored fins lifting off, flame underneath, and a small tower at the lower left. A line under the rocket reads “Running on NVIDIA AI Infrastructure.” Seventeen divided by 6.4 is about 2.7. That figure is arithmetic on the two labels in the frame. The blog’s written claim is up to eight times faster token generation than Astra Standard. The frame does not print the words eight times, and it does not print a calendar date. It is the comparison graphic. It is not a photograph of a data center.
In plain terms, NVIDIA said on Thursday that GPT-6 Astra Ultrafast is available now in the OpenAI API and to eligible ChatGPT Work and Codex users, running on Blackwell GPUs. NVIDIA says the mode can generate tokens up to eight times faster than Astra Standard, for coding loops, tool calls, and interactive apps. Philippe Tillet, OpenAI’s inference lead, said NVIDIA’s tooling helped OpenAI’s models program Blackwell and Rubin GPUs, and that Astra turns that into fast kernels across wait, throughput, and cost. Uday Ruddarraju, OpenAI’s compute chief technology officer, said OpenAI used its own models to optimize inference on NVIDIA GPUs. NVIDIA says that work continues after deployment, and it points developers to OpenAI’s Ultrafast guide for access, pricing, and how to switch the mode on. The guide says to use the tier when the speed is worth the higher cost. It does not, on the lines used here, print a dollar price.
RELATED
Sources
- NVIDIA Blog — GPT-6 Astra Ultrafast on Blackwell, 1 Oct 2026
blogs.nvidia.com
- OpenAI — Ultrafast mode guide
developers.openai.com