The AI agent market has a new leader. SpaceXAI's Grok 4.5 model has not only surpassed Anthropic's flagships — Claude Fable 5 and Claude Opus 4.8 — in the independent AutomationBench-AA benchmark, but it has done so with a massive lead in cost efficiency. This is not just a test victory; it's a paradigm shift.

Grok 4.5, built on the V9 architecture with 1.5 trillion parameters, achieved a score of 51.4% in Artificial Analysis's AutomationBench-AA benchmark. For comparison, Claude Fable 5 scored 48.6%, and Claude Opus 4.8 scored 48.5%. However, the key metric is not just accuracy, but also cost.

Grok 4.5 completes a task for $0.34, while Fable 5 costs $1.35 and Opus 4.8 costs $1.46. The closest competitor in price, Gemini 3.5 Flash, costs $0.49 per task. This low cost is achieved through radical efficiency: Grok 4.5 uses about 8,000 output tokens per task — just 25% of the resources consumed by Opus 4.8. The total token consumption of 0.44 million per task is one of the lowest on the market.

Financial Sector: Where Grok 4.5 Truly Shone

The model's most telling result is in the financial category — the most difficult in the test. Grok 4.5 scored 71%, leaving Fable 5 (64%) and Opus 4.8 (62%) behind. For companies automating financial processes, this gap is critical. An error in a financial agent can lead to direct losses, and here Grok demonstrates clear superiority.

However, there is a nuance: Grok 4.5 violates rules an average of 0.63 times per task, which is higher than Opus 4.8 (0.55) and Gemini 3.5 Flash (0.46). This is the price for aggressive optimization of speed and cost. For businesses deploying agents in production environments, this factor requires special attention and additional validation layers.

The test includes 657 tasks across 40 simulated applications, including Gmail, Slack, Salesforce, and HubSpot. The final score reflects the proportion of tasks completed by the agent without violating the established rules. Artificial Analysis keeps the task list confidential, ensuring objective results and preventing model "training to the test."

My professional perspective: Grok 4.5 sets a new standard for the AI agent market. The combination of high accuracy and radically low cost is exactly what is needed for the mass adoption of agents in business processes. The victory in the financial sector is a signal for the entire FinTech industry. If Anthropic and Google do not respond with adequate price reductions, Grok's share of the corporate market will grow exponentially.