OpenAI has optimized the default response speed of GPT-6 Astra and GPT-6.1 Sol, with OpenAI’s Codex engineering lead saying the models should feel roughly 50% faster for subscription users in supported products and Sign in with ChatGPT partner apps. The figure is OpenAI’s claim, not a measured guarantee for every task. The announced default rate moves from roughly 30 to roughly 50 output tokens per second, with no settings change, according to Sottiaux’s October 5 first-party post.

The improvement is notable because OpenAI is competing not only on model intelligence but increasingly on how quickly AI systems can respond and complete work. Faster token generation can make coding, writing and agentic workflows feel substantially more responsive, although raw tokens-per-second should not be confused with overall task-completion speed. OpenAI’s official documentation also lists separate Fast and Ultrafast modes for its latest models, showing that inference speed has become an increasingly explicit part of the company’s product strategy.

Key takeaways

  • GPT-6 Astra and GPT-6.1 Sol are reportedly getting about a 50% default speed improvement.
  • OpenAI’s Thibault Sottiaux said generation increased from roughly 30 to roughly 50 tokens per second.
  • The optimization applies across OpenAI’s products and supported Sign in with ChatGPT partner applications.
  • OpenCode, Pi, Amp and Devin were specifically mentioned among supported partners.
  • The speed improvement does not mean the models became 50% more intelligent.
  • A 30-to-50 token-per-second change is mathematically closer to a 67% increase, suggesting the published numbers are approximate.
  • Independent Artificial Analysis measurements currently put GPT-6.1 Sol High at about 50.4 tokens per second and GPT-6 Astra High at about 47.7 tokens per second.
  • OpenAI is separately offering faster inference modes, including Ultrafast, which can reach substantially higher generation rates than the standard experience.

OpenAI speeds up its latest GPT-6 models

OpenAI has begun rolling out a default inference optimization for two of its latest models: GPT-6 Astra and GPT-6.1 Sol.

Thibault Sottiaux, OpenAI’s Codex engineering lead, announced the change on October 5 as the first improvement in a 28-day shipping initiative, as The New Stack reported. He said the models had been optimized to be approximately 50% faster by default across OpenAI products and partner applications using Sign in with ChatGPT.

The change is designed to happen automatically.

Users do not need to select a faster mode or change a setting to receive the default optimization.

Sottiaux subsequently put a more concrete number on the improvement, saying the models could now generate approximately 50 tokens per second, compared with approximately 30 tokens per second previously.

That distinction matters because token-generation speed is only one component of AI performance.

A model can generate tokens quickly but still take a long time to complete a task if it needs to reason extensively, call tools, browse websites, execute code or wait for external systems.

From 30 to roughly 50 tokens per second

The headline number is straightforward:

Before: ~30 tokens/second
After: ~50 tokens/second

If both numbers are treated as exact, the mathematical increase is approximately 67%, not 50%.

That does not necessarily mean OpenAI’s announcement is contradictory.

The figures appear to be approximate, and the company’s statement describes the overall optimization as roughly 50% faster rather than publishing a formal controlled benchmark showing an exact 66.7% increase.

That distinction is important for reporting.

The safest interpretation is that OpenAI has made the models materially faster, with the company describing the improvement as around 50% and giving approximate generation rates of 30 and 50 tokens per second.

It should not be presented as a precise laboratory measurement.

OpenAI claimed default output speed comparisonOpenAI said the approximate default output rate rose from 30 to 50 tokens per second for GPT-6 Astra and GPT-6.1 Sol in supported subscription experiences. These are company figures, not independently verified before-and-after results.Default output rate claimed by OpenAIBefore~30 tokens/sAfter~50 tokens/sSource: OpenAI staff statement, 5 October 2026; values rounded

What does 50 tokens per second actually mean?

A token is a small unit of text processed by an AI model.

Tokens do not correspond exactly to words.

For English, a token may be part of a word, an entire short word or punctuation. As a rough illustration, a response containing several hundred words can contain considerably more tokens.

At 50 tokens per second, an AI model can stream text noticeably faster than one producing 30 tokens per second.

The difference is especially visible in long responses.

For example, if a model needs to generate 1,000 output tokens and everything else remains constant:

  • At 30 tokens/second: about 33 seconds of generation
  • At 50 tokens/second: about 20 seconds of generation

That is a theoretical comparison, not a prediction of real-world ChatGPT response times.

Actual latency can include time spent waiting for the first token, model reasoning, tool calls, network communication and other processing.

This is why OpenAI and independent benchmarking organizations distinguish between output speed, time to first token and end-to-end task time.

Independent measurements are already around 50 tokens per second

There is some useful independent context for OpenAI’s claim.

Artificial Analysis’s separate API measurements put GPT-6.1 Sol at approximately 50.4 tokens per second at high reasoning effort, compared with approximately 47.7 tokens per second for GPT-6 Astra. These figures can change and do not establish the before-and-after subscription improvement.

Those measurements are broadly consistent with the roughly 50-token-per-second figure being discussed.

However, they should not be interpreted as proof of the exact before-and-after improvement described by OpenAI.

Artificial Analysis is measuring a particular model configuration and provider environment, while Sottiaux’s announcement concerns a default optimization across OpenAI products and partner applications.

Different serving environments can produce different speeds.

GPT-6 Astra remains OpenAI’s flagship model

GPT-6 Astra is OpenAI’s highest-end model in the GPT-6 family.

OpenAI describes Astra as its most intelligent model for demanding reasoning, coding, computer use, science and professional work. The company says Astra achieved state-of-the-art results across several of its evaluations, including computer-use and scientific benchmarks.

The model is also designed for agentic workflows.

It can use computers, browse, work with code and execute multi-step professional tasks.

That makes speed particularly important.

For a conventional chatbot, reducing the time required to stream a response is useful.

For an AI agent, the effect can compound.

An agent may need to:

  1. Understand a task.
  2. Plan the next action.
  3. Generate an instruction.
  4. Call a tool.
  5. Read the result.
  6. Reason about the result.
  7. Generate another action.
  8. Repeat the process.

If each model interaction becomes faster, the total workflow can potentially become more responsive.

But again, model output speed is only one part of that equation.

GPT-6.1 Sol is positioned as the faster-value option

GPT-6.1 Sol occupies a different position in OpenAI’s model lineup.

OpenAI describes it as offering near-Astra performance at substantially lower cost.

OpenAI’s GPT-6.1 Sol announcement lists it at $2 per million input tokens and $10 per million output tokens, compared with $10 per million input tokens and $50 per million output tokens for GPT-6 Astra.

That means Sol is positioned for organizations and developers that want high-end capabilities without paying Astra-level API prices.

The speed improvement therefore has a second implication.

If Sol can become both cheaper and faster while remaining close to Astra on important workloads, it could become a more attractive default model for production AI applications.

OpenAI is explicitly positioning Sol as the balance between intelligence, cost and speed.

OpenAI has been pushing speed as a product feature

The latest optimization is not happening in isolation.

OpenAI has increasingly separated its model capabilities from the infrastructure used to serve them.

The company now offers Fast mode for supported models and has introduced Ultrafast for eligible users and workloads.

OpenAI says Fast mode can increase speed for supported models, including GPT-6 Astra and GPT-6.1 Sol.

For GPT-6 Astra, OpenAI also offers an Ultrafast mode.

The company’s documentation says Astra can generate tokens at up to 8 times the standard speed in Codex under Ultrafast, while its API documentation describes a separate high-speed inference configuration.

VentureBeat reported that OpenAI’s new Ultrafast tier can reach up to 300 tokens per second, depending on the environment and configuration.

This means the approximately 50-token-per-second default should not be confused with the maximum generation speed available through OpenAI’s premium inference options.

Why faster AI matters for coding

Coding is one of the areas where users are likely to notice a speed improvement.

AI coding agents often generate large amounts of code, inspect files and iterate through multiple steps.

A faster model can reduce the waiting time between individual stages.

For example, a coding workflow might involve:

Prompt → reasoning → code → execution → error → reasoning → correction → test

Even if only the text-generation portion becomes faster, the entire workflow can feel more responsive.

GPT-6 Astra is already positioned by OpenAI as its strongest model for software engineering. OpenAI’s published evaluations include coding-agent benchmarks such as Terminal-Bench and DeepSWE.

GPT-6.1 Sol has also improved substantially over the previous Sol model on coding evaluations.

OpenAI says GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1 while doing so at approximately one-fifth of Astra’s cost under the tested conditions.

Adding faster inference makes that combination more attractive for developers.

Faster tokens do not automatically mean faster agents

There is an important limitation to the speed story.

A model generating 50 tokens per second does not necessarily complete a task 67% faster.

Agentic tasks can contain many components that have nothing to do with token generation.

Consider an AI coding agent that spends:

  • 10 seconds planning
  • 20 seconds generating code
  • 30 seconds running tests
  • 15 seconds reading files
  • 10 seconds waiting for a tool

If generation becomes dramatically faster, only the generation component improves.

The other 55 seconds remain.

This is why end-to-end task completion is arguably a more useful measure of AI-agent performance than raw token throughput.

OpenAI itself highlights task-level measurements in its GPT-6.1 Sol announcement, including computer use, professional workflows and scientific research rather than relying exclusively on generation speed.

OpenAI is also optimizing token efficiency

There is another factor behind the latest generation of model improvements: models can become more efficient not only by generating tokens faster, but also by needing fewer tokens to accomplish the same task.

OpenAI has emphasized that GPT-6 Astra can achieve strong performance while using fewer output tokens in certain workflows. Its GPT-6.1 Sol announcement similarly highlights cost and task efficiency across several evaluations.

This creates two different paths toward faster AI:

Path 1: Generate tokens faster

Path 2: Need fewer tokens

The most useful improvements combine both.

A model that generates 50 tokens per second but requires 20,000 tokens to complete a task may still be slower than a model that generates 35 tokens per second but solves the same task in 8,000 tokens.

That is increasingly relevant as AI moves from simple question answering toward long-running agents.

The speedup extends beyond OpenAI’s own apps

One of the more interesting elements of the announcement is its reach.

Sottiaux said the optimization applies across OpenAI products and partners using Sign in with ChatGPT.

Sottiaux specifically named OpenCode, Pi, Amp and Devin among the partner applications affected, though availability depends on each partner’s support for Sign in with ChatGPT.

That means OpenAI is not treating inference speed purely as a ChatGPT feature.

It is also becoming part of the broader ecosystem built around OpenAI accounts and models.

For developers, this could make OpenAI’s models more attractive as the backend for third-party coding and productivity applications.

For OpenAI, it also strengthens the value of its subscription ecosystem.

A user can potentially access a faster model experience across several products without changing providers.

Scope of the claimed speed updateA subscription user selects GPT-6 Astra or GPT-6.1 Sol in a supported OpenAI product or Sign in with ChatGPT partner. The announcement does not establish a faster OpenAI API or universal availability in every chat.Subscription usereligible model accessGPT-6 Astra or Soldefault serving pathSupported OpenAI orpartner appWhere the announcement appliesAPI speed and availability in standard ChatGPT chat were not announced.

The timing is significant for OpenAI

The announcement comes shortly after the release of GPT-6.1 Sol.

OpenAI introduced Sol as a more capable version of GPT-6 Sol that approaches Astra’s performance on several professional tasks while costing substantially less.

The company is therefore simultaneously improving three dimensions:

Capability → GPT-6.1 Sol

Cost → lower token prices

Speed → faster inference

That combination is strategically important.

The AI market is no longer competing solely on benchmark scores.

Users increasingly care about how quickly a model responds, how much it costs and how reliably it completes a workflow.

A slightly less capable model that is substantially cheaper and faster can sometimes be more useful in production than a theoretically stronger model.

The rise of real-time AI

The broader trend is toward AI that feels less like submitting a request to a remote computer and more like interacting with a responsive assistant.

This matters particularly for:

  • Coding assistants
  • Voice applications
  • Computer-use agents
  • Customer-support agents
  • Interactive education
  • Real-time research
  • AI-powered development environments
  • Autonomous software agents

In these settings, latency affects user experience directly.

A delay of several seconds may be acceptable when generating a long research report.

The same delay can feel frustrating when an AI assistant is supposed to collaborate interactively.

OpenAI’s investments in Fast and Ultrafast modes indicate that the company increasingly sees latency as a differentiating feature rather than simply an infrastructure detail.

What users should expect

For eligible subscribers using GPT-6 Astra or GPT-6.1 Sol in OpenAI products or Sign in with ChatGPT partner apps, the expected effect is faster output streaming under the updated default serving configuration. OpenAI has not said this makes those models available in every ordinary ChatGPT chat.

Users do not need to activate a new model setting for the reported default optimization.

However, the actual experience can vary.

Network conditions, model reasoning effort, system load, prompt complexity, tool calls and the particular application can all affect perceived latency.

A user should therefore not expect every response to become exactly 50% faster.

The reported 50-token-per-second figure is also an approximate generation rate, not a promise that every answer will appear at that speed from its first token to its last.

What this means for AI competition

The speed race is becoming another major battleground between OpenAI, Google, Anthropic and specialized inference providers.

Model developers are increasingly looking for ways to deliver frontier-level intelligence without forcing users to wait for every reasoning step.

That requires improvements throughout the stack:

  • Better model architectures
  • More efficient inference
  • Optimized serving infrastructure
  • Better tokenization
  • Faster hardware
  • Smarter caching
  • Speculative decoding
  • Distillation
  • Improved agent orchestration

The result is a shift from asking “How intelligent is the model?” toward a broader question:

“How much useful work can the model complete per second and per dollar?”

That may become one of the most important metrics for the next phase of AI competition.

The Bigger Picture

OpenAI’s latest GPT-6 speed improvement illustrates how AI competition is moving beyond model launches. GPT-6 Astra and GPT-6.1 Sol are already positioned around different combinations of intelligence and cost, while Fast and Ultrafast modes add another layer of performance optimization.

The reported move toward roughly 50 tokens per second is meaningful because faster inference can make AI agents feel more interactive and reduce waiting during long coding or professional workflows. But raw token speed should not become the industry’s only measure of progress. End-to-end completion time, reliability, tool latency, reasoning efficiency and cost per completed task will matter more as AI systems increasingly perform work rather than simply generate text.

Looking Ahead

OpenAI’s next challenge will be turning faster inference into faster completed work. If the company can combine higher token throughput with shorter reasoning paths, better tool execution and stronger agent orchestration, users could see much larger improvements in practical productivity than the headline generation-rate increase suggests.

The longer-term competition will likely be defined by a combination of intelligence, speed and economics. A model that is slightly less capable but dramatically cheaper and faster can win large production workloads, while flagship systems such as Astra can remain focused on the hardest tasks. For OpenAI, making both ends of the GPT-6 family faster strengthens its position in a market where AI users increasingly expect models to respond almost as quickly as conventional software.

Sources and related coverage

The October 5 speed claim comes from Sottiaux’s original October 5 post. Independent coverage by The New Stack, ADSLZone and Crypto Briefing separately reports the announcement; none independently proves a universal 50% user-experience gain. For model context, see our report on GPT-6 Sol and Luna pricing and our report on OpenAI’s separate Ultrafast mode.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.