The AI agent market has a new favorite. Results from the independent AutomationBench-AA benchmark by Artificial Analysis showed that SpaceXAI's Grok 4.5 model has not only caught up but confidently surpassed heavyweights like Claude Fable 5 and Claude Opus 4.8. Elon Musk, commenting on the launch, emphasized speed and cost efficiency — and the numbers fully confirm his words.
The key metric is the final score. Grok 4.5 achieved 51.4%, leaving Fable 5 at 48.6% and Opus 4.8 at 48.5% behind. However, what's far more interesting is the cost. Completing one task costs just $0.34. For comparison, Fable 5 costs $1.35, and Opus 4.8 costs $1.46. The closest competitor in price is Gemini 3.5 Flash ($0.49), but it lags in quality.
Efficiency as the Main Trump Card
The secret behind such a low price lies in the V9 architecture with 1.5 trillion parameters. Grok 4.5 uses about 8,000 output tokens per task — just 25% of the resources consumed by Opus 4.8. The total consumption of 0.44 million tokens per task is one of the lowest on the market. It is this efficiency and low token price that create the decisive advantage.
Beyond the overall statistics, details matter. Grok 4.5 correctly completed 79.9% of individual actions, and fully, from start to finish, completed 21.9% of tasks. In the financial category — the most challenging in the test — the model achieved 71% successful solutions. This is significantly higher than Fable 5 (64%) and Opus 4.8 (62%).
However, there is a nuance. Grok 4.5 violates rules more often than competitors: an average of 0.63 violations per task compared to 0.55 for Opus 4.8 and 0.46 for Gemini 3.5 Flash. For companies deploying agents in real financial systems, this is a critical point — one wrong action could lead to serious losses.
Analyst's opinion: Grok 4.5 sets a new standard for the AI agent market, shifting the focus from "pure" performance to economic efficiency. However, the increased tendency for errors is a "time bomb." For now, the model wins due to its low price, but for the enterprise sector, reliability remains a priority. The next version will need to address the rule-compliance issue to solidify its success.