Crypto news

14.08.2026
08:10

Speed race: Google and OpenAI unveiled next-generation AI solutions

битва роботов, AI battle

On August 13, a double blow hit the artificial intelligence market: Google announced the Gemini 3.7 Flash model, while OpenAI introduced a revolutionary Ultrafast mode for GPT-5.6 Sol. Both developments target a critically important segment—high-speed execution of tasks related to programming and autonomous AI agents.

Gemini 3.7 Flash: A Bet on Efficiency and Cost Savings

Google's new model is positioned as a workhorse for the corporate sector and developers. Key improvements concern multi-stage planning, precise instruction following, and code generation. However, the most intriguing aspect has been the radical price reduction: the cost of use has dropped by half relative to the previous version's pricing.

The Flash line traditionally covers scenarios where latency is critical: coding, agent pipelines, and tasks with multiple sequential model calls. Recall that in July, Gemini 3.6 Flash was released with a price tag of $1.5 per million input tokens and $7.5 per million output tokens. At that time, the company also reported a 17% reduction in output token consumption compared to 3.5 Flash.

Notably, the flagship Gemini 3.5 Pro is still in closed testing with partners, and the timeline for its public release remains unclear. This creates a curious imbalance: Google is clearly betting on the mass-market segment rather than premium models.

OpenAI Ultrafast: 750 Tokens Per Second on Cerebras

OpenAI responded with its own trump card—the Ultrafast mode for GPT-5.6 Sol. The claimed figures are impressive: up to 14x acceleration relative to the standard version and generation of up to 750 output tokens per second. The technological foundation is Cerebras infrastructure, which specializes in giant silicon wafers.

At launch, only a limited pool of clients will get access to Ultrafast via the API. This is a deliberate step: OpenAI will gradually scale up computing capacity, connecting new users incrementally. Notably, the figure of 750 tokens per second had already appeared in the June announcement of GPT-5.6, but at that time the company warned of uneven access to the accelerated version.

The context of the race is amplified by the August release of Grok 4.6 from SpaceXAI, focused on long-term agent scenarios and interactive applications.

My analysis: We are witnessing a tectonic shift in the competitive struggle—from answer quality to the speed of obtaining them. The cheapening of Google's Flash models and OpenAI's aggressive ultra-fast mode are a direct consequence of pressure from open models. However, 750 tokens per second is not just marketing: such figures open the door for real-time agents that can interact with users without noticeable delay. The only question is whether Cerebras can scale quickly enough to democratize this technology, or whether ultra-speed will remain a privilege of select corporate clients.