Crypto news

14.08.2026
07:28

Race for Speed: Google and OpenAI Unveiled Next-Generation AI Solutions

битва роботов, AI battle

The artificial intelligence market continues to be in turmoil: on August 13, two tech giants simultaneously rolled out updates targeting the same niche—ultra-fast task processing for developers and AI agents. Google introduced Gemini 3.7 Flash, while OpenAI announced an Ultrafast mode for its flagship model GPT-5.6 Sol. This is not just evolution, but a clear signal of where the industry is heading: toward speed and efficiency, not just raw power.

What's new in Gemini 3.7 Flash

Google positions Gemini 3.7 Flash as a workhorse for programming and business process automation. Key improvements focus on multi-step planning, precise instruction following, and code generation. However, the most interesting part is the price: the company announced a halving of usage costs compared to the starting price of the previous version. This is an aggressive move, given that the Flash line was already aimed at scenarios where response speed is critical: coding, agent work, and tasks with many sequential calls to the model.

As a reminder, in July, Gemini 3.6 Flash was released with a price of $1.5 per 1 million input tokens and $7.5 per 1 million output tokens. At the same time, the company reported that the new model consumes 17% fewer output tokens than 3.5 Flash. Meanwhile, the flagship Gemini 3.5 Pro, which many were waiting for, has still not been released to the public—developers are still testing it with partners, and the release timeline remains unclear.

How OpenAI accelerated GPT-5.6 Sol

OpenAI responded symmetrically, but with a different emphasis. The Ultrafast mode for GPT-5.6 Sol, according to the company, speeds up generation to 750 output tokens per second—up to 14 times faster than the standard mode. The secret lies in using Cerebras infrastructure, which specializes in ultra-high-speed computing. However, at launch, only a limited number of clients will get access to Ultrafast via the API, and OpenAI promises to expand access as computing capacity grows.

The figure of 750 tokens per second is not new—it was mentioned back in June during the GPT-5.6 announcement. It was clear then that Sol would run on Cerebras, but with a caveat about phased user onboarding. Now that plan is starting to materialize. Against this backdrop, it's worth recalling the August release of Grok 4.6 from SpaceXAI—a model focused on long-term agentic tasks and programming. Competition is heating up to the limit.

My analysis: Both announcements confirm the trend toward "small and fast" models that can operate in real time. Google's price cut and OpenAI's extreme speed are two different approaches to the same problem: making AI accessible for mass automation. In the short term, the winner will be the one who can offer the best balance of price and response latency, especially in the agentic solutions segment, where every extra millisecond is lost money.