Google and OpenAI are accelerating the AI race: lightning-fast models for code and agents announced

Competition in the artificial intelligence market has reached a new level of speed. On August 13, two tech giants unveiled their ultra-fast solutions: Google introduced the Gemini 3.7 Flash model, while OpenAI announced the Ultrafast mode for GPT-5.6 Sol. Both developments target the segment where response time is critical—programming and autonomous AI agent operations.
Gemini 3.7 Flash: A Focus on Efficiency and Cost Savings
Google positions the new release as a tool for automating business processes and coding. Key improvements relate to multi-step planning and instruction following, which directly impact code generation quality. Notably, the cost of usage has been halved relative to the starting price of the previous version, making the model more accessible for mass integration.
The Flash line traditionally covers scenarios requiring not just accurate but also instant responses: this includes AI agent work and tasks with multiple sequential calls. For context, in July, Gemini 3.6 Flash was released with a price tag of $1.5 per million input tokens and $7.5 per million output tokens. At that time, a 17% savings on output tokens compared to 3.5 Flash was claimed. However, the flagship Gemini 3.5 Pro, which many were waiting for, never appeared—testing with partners is ongoing, and the exact release timeline remains unclear.
OpenAI Ultrafast: 750 Tokens Per Second Powered by Cerebras
OpenAI has gone even further in radical acceleration. The new Ultrafast mode for GPT-5.6 Sol can deliver up to 750 output tokens per second, which, according to the company, is 14 times faster than the standard version. The technological foundation is Cerebras infrastructure, which specializes in ultra-high-speed computing.
At launch, the mode is available to a limited set of clients via API, and OpenAI will gradually onboard users as capacity expands. The 750-token figure was mentioned back in June during the GPT-5.6 announcement, but developers honestly warned then that not everyone would get instant access to the accelerated version. In this same context, it is worth noting that in August, the SpaceXAI team released Grok 4.6, betting on long-term agentic tasks and interactive applications.
My conclusion: We are witnessing a clear trend: the battle is shifting from the "smarter" dimension to the "faster and cheaper" dimension. For Web3 and the crypto industry, this is critically important, as AI agents are becoming a key tool for automating DeFi operations and market analysis. The price reduction on Gemini Flash is a direct signal for developers: infrastructure costs for AI will decline, opening up new opportunities for scaling projects.