OpenAI has introduced Ultrafast, a new service tier for GPT-5.6 Sol that it says can run up to 14× faster than Standard processing. The feature launches first via the OpenAI API and is currently available in limited preview for select customers.
Powered by Cerebras infrastructure, Ultrafast can generate up to 750 output tokens per second. In simple terms, this means the model can produce long responses, analyze large inputs, and complete multi-step reasoning tasks with much lower waiting time.
OpenAI is positioning the service for products where response delays can affect business operations, customer experience, or security decisions. The launch is significant because high-speed AI services have often required users to select a smaller or less capable model.
OpenAI says Ultrafast is designed to bring the intelligence of GPT-5.6 Sol to real-time workflows without that trade-off. The company describes this direction as delivering more useful work per second, rather than simply making responses appear faster.
Ultrafast could be particularly useful in cybersecurity incident response. During an active outage or suspected compromise, defenders must quickly review logs, alerts, traces, recent code changes, and internal communications.
A faster model could help analysts correlate evidence, identify likely causes, recommend validation checks, and prepare remediation steps while an incident is still developing.
OpenAI Unveils Ultrafast Mode in GPT‑5.6 Sol
For example, an operations team responding to suspicious activity could feed the model authentication logs, endpoint telemetry, cloud audit records, and a timeline of recent deployment changes.
Instead of waiting for a long analysis, the team could receive a rapid summary of anomalous activity and a prioritized set of investigation paths. Human analysts would still need to validate findings and approve containment or deployment actions.
OpenAI also highlighted financial security and fraud analysis as potential use cases. Organizations could use the higher-speed tier to assess changing transaction patterns, investigate suspicious behavior, and support analysts during time-sensitive events.
These workflows need careful controls because fast model output is not the same as verified evidence. OpenAI identified customer support, voice applications, commerce, coding, research, and experimentation as early use cases.
In customer support, low latency could allow an AI assistant to consult multiple internal systems and answer complex questions during a live conversation.
In commerce, it could answer product questions, check inventory, suggest personalized recommendations, and resolve checkout issues before a shopper leaves the site.
For research teams, Ultrafast may shorten the cycle from hypothesis to experiment, result review, and follow-up testing. OpenAI said internal teams are exploring whether workloads previously handled as overnight batch jobs can instead be completed interactively during the day.
Cerebras is powering the low-latency inference behind Ultrafast. The companies say the tier maintains the same GPT-5.6 Sol intelligence while substantially increasing output speed over OpenAI’s Standard processing tier.
Access remains restricted during the preview phase. OpenAI is using the early deployment to evaluate which business workflows gain the most value from the speed increase, and says access will expand as capacity becomes available.
Strengthen Your SOC by Accelerating Threat Detection & Rapid Investigations. -> Integrate ANY.RUN With Your SOC Now.

