OpenAI Previews Ultrafast Mode: GPT-5.6 Sol at 14x Speed

OpenAI is previewing a new Ultrafast mode that can run GPT-5.6 Sol up to 14 times faster than standard processing, generating up to 750 output tokens per second. The mode launches first through the OpenAI API and is powered by Cerebras hardware, with the company targeting latency-sensitive production workloads where speed matters as much as model capability.

ChatGPT Ultrafast Mode

The use cases OpenAI is pointing to include voice interfaces, customer support, commerce, developer agents, financial research, and security response. OpenAI says its own developers have used Ultrafast mode to analyze logs and traces during incidents and to compress research cycles that previously ran overnight into multiple runs within the same window. At 750 tokens per second, a 1,000-token response completes in well under two seconds, a threshold that matters for anything operating inside a user-facing product.

GPT-5.6 Sol’s Position in the Model Lineup

Sol is OpenAI’s strongest model in the GPT-5.6 family, which the company introduced in June and made broadly available in July through ChatGPT, Codex, and the API. The family includes three tiers:

  • Sol, flagship model with agentic improvements in coding, biology, and cybersecurity; priced at $5 input / $30 output per 1M tokens
  • Terra, balanced everyday model, similar performance to GPT-5.5 but 2x cheaper; priced at $2.50 input / $15 output per 1M tokens
  • Luna, fast and affordable, OpenAI’s lowest-price option; priced at $1 input / $6 output per 1M tokens

The GPT-5.6 rollout followed a limited preview period triggered by U.S. Government review, with the Trump administration requesting oversight of frontier AI models before launch. Sol also includes a new “max” reasoning effort setting and an “ultra” mode that uses sub-agents for complex work, along with what OpenAI describes as its most safety stack to date.

Ultrafast Comes After Recent Accuracy Improvements

The Ultrafast mode announcement follows a separate update earlier this month in which OpenAI said it is tuning GPT-5.6 Sol specifically for everyday ChatGPT conversations, aiming for more focused answers, appropriate detail levels, and less unnecessary formatting. In an internal evaluation covering financial, medical, and legal prompts, answers containing at least one factual error were reportedly 68% less common with GPT-5.6 Sol than with GPT-5.5 Instant. That accuracy work and the Ultrafast mode together suggest OpenAI is trying to close the gap on two fronts at once – making the model more reliable for high-stakes domains while also making it fast enough for real-time applications that previously required smaller, less capable models.

About the Author

Imran Hussain is the founder and editor of iThinkDifferent, which he launched in 2008 to cover Apple news, reviews, and how-to guides. He has spent over 15 years writing about iOS, macOS, and the wider Apple ecosystem, with a focus on hands-on guides - installing developer betas, troubleshooting, and walking through new features on his own devices. Based in Dubai, he also loves to cover photography, gaming, and the tech industry more broadly on his social media profiles.

Leave a Reply