The AI speed race: Google and OpenAI unveiled new solutions for developers

On August 13, a double announcement took place that set a new direction in the development of generative AI. Google introduced the Gemini 3.7 Flash model, while OpenAI announced the Ultrafast mode for its GPT-5.6 Sol. Both companies are clearly betting on speed and efficiency, which is critical for tasks involving code and autonomous AI agents.
What's new in Gemini 3.7 Flash
Google positions Gemini 3.7 Flash as a workhorse for programming and business process automation. Developers claim significant improvements in multi-step planning, instruction-following accuracy, and code generation quality. At the same time, the company has taken a radical step: the cost of usage has been halved compared to the starting price of the previous version. This makes the model extremely attractive for large-scale production workloads.

The Flash line was originally conceived as a solution for scenarios where response speed is no less important than quality. This includes coding, AI agent work, and any tasks involving many sequential calls to the model. Recall that in July, Gemini 3.6 Flash was released at a price of $1.5 per 1 million input tokens and $7.5 per 1 million output tokens. At that time, the company also reported that the new model consumes 17% fewer output tokens than its predecessor.
However, in the shadow of these announcements remains the question of the flagship Gemini 3.5 Pro. The model has still not been publicly released, and developers are only testing it with select partners. Specific release timelines remain vague, which creates intrigue in the market.
How OpenAI accelerated GPT-5.6 Sol
OpenAI, for its part, introduced the Ultrafast mode for GPT-5.6 Sol. The stated figures are impressive: operating speeds up to 14 times higher than standard and generation of up to 750 output tokens per second. The secret to this acceleration lies in the use of Cerebras infrastructure, which specializes in ultra-fast AI chips.
At launch, Ultrafast will be available via API to a limited group of clients. This is a logical step: OpenAI needs to test the load and scale up computing power. As it grows, access will be opened to a wider audience. The figure of 750 tokens per second was already mentioned in June when GPT-5.6 was announced. At that time, developers warned that the accelerated version on Cerebras would be rolled out in stages.
In this context, it's worth recalling the August release of Grok 4.6 from SpaceXAI, which focused on long-term agentic tasks and programming. Competition in the "fast" models segment is heating up, and that's great news for the market.
My analysis: Both announcements confirm the key trend of 2026 — a shift in focus from "smart" models to "fast and cheap" ones. For Web3 developers and DeFi protocols, this means lower operational costs and the ability to build more complex agents operating in real time. Halving the price of Gemini 3.7 Flash is not just marketing, but a direct attempt to capture market share in enterprise development, where the price per token is often a decisive factor.