Speed race: Google and OpenAI unveiled next-generation AI solutions for coding and agents

On August 13, tech giants almost simultaneously updated their AI product lines. Google introduced Gemini 3.7 Flash, while OpenAI announced the Ultrafast mode for GPT-5.6 Sol. Both developments target the same segment—high-speed execution of tasks related to code generation and AI agent operations. This is not just a coincidence but a clear signal to the market: the battle for performance and latency is entering a new level.
What's new in Gemini 3.7 Flash
Google positions Gemini 3.7 Flash as a workhorse for programming and business process automation. The company claims significant improvements in multi-step planning, instruction following, and code generation quality. Notably, the cost of using the model has halved compared to the launch price of the previous version.
The Flash line has historically been designed for scenarios where response speed is critical, not just its depth. This includes coding, autonomous agent work, and tasks with many sequential calls to the neural network. In July, Google released Gemini 3.6 Flash at $1.5 per 1 million input tokens and $7.5 per 1 million output tokens. That model showed 17% greater output token efficiency than 3.5 Flash.
However, the flagship Gemini 3.5 Pro has still not been released to the public. Developers are currently testing it with a limited circle of partners, and the exact release timeline remains unknown. This creates intrigue: while Google focuses on the speed of budget models, the premium segment remains without an update.
How OpenAI accelerated GPT-5.6 Sol
OpenAI, for its part, took the path of hardware acceleration. The new Ultrafast mode for GPT-5.6 Sol can operate up to 14 times faster than the standard version, delivering up to 750 output tokens per second. This is achieved through the use of Cerebras infrastructure—specialized chips optimized for large language models.
At launch, Ultrafast is available via API only to a limited number of clients. OpenAI plans to gradually expand access as computing capacity grows. The figure of 750 tokens per second was already mentioned by the company during the GPT-5.6 announcement in June, with a warning that the accelerated version would initially not be available to everyone.
Recall that in August, SpaceXAI also introduced Grok 4.6—a flagship model focused on long-term agent tasks and programming. Competition in this segment is becoming increasingly fierce.
My analysis: We are witnessing a fundamental shift in the industry—competition is moving from simple answer quality to inference speed and cost. Halving the price of Gemini 3.7 Flash is a direct response to pressure from open models, while the race for tokens per second is an attempt to capture the real-time market, where AI agents will execute transactions and write code live. In the coming months, we can expect even more aggressive price dumping on APIs.