SpaceXAI has officially launched Grok 4.5 — a new flagship model focused on programming, agent work, and solving complex intellectual tasks. In the press release, the model is described as the "most powerful system" in the company's portfolio. However, the key emphasis is not on absolute superiority in benchmarks, but on the balance of price, speed, and efficiency.
Grok 4.5 is already available via the API and is also integrated into the Grok Build and Cursor platforms across all pricing plans. Users in the European Union are still left out — the launch in the region is expected in mid-July.
Pricing as the Main Trump Card
SpaceXAI offers Grok 4.5 at a price of $2 per 1 million input tokens and $6 per 1 million output tokens. This is significantly cheaper than direct competitors. For comparison: Claude Opus 4.8 from Anthropic costs $5 and $25 respectively, Claude Fable 5 — $10 and $50, and GPT-5.6 Sol from OpenAI — $5 and $30. Even within OpenAI's own lineup, the Luna model is offered at $1 for input tokens, but $6 for output tokens, which is comparable to Grok 4.5.
Elon Musk described the new model as "an Opus-class model, but faster, cheaper, and more token-efficient." According to him, internal SpaceXAI tests show that Grok 4.5 is "roughly comparable to Opus 4.7, but much faster."
Benchmark Results: Mixed, but Promising
The data published by SpaceXAI presents a mixed picture. On the DeepSWE 1.0 test, Grok 4.5 scored 62%, trailing Fable (66.1%) and GPT-5.5 (64.31%), but surpassing Claude Opus 4.8 (55.75%) and Opus 4.7 (40.12%). The newer DeepSWE 1.1 benchmark showed a result of 53% — lower than Opus 4.8 (59%), GPT-5.5 (67%), and Fable (70%).
On SWE Bench Pro, the model scored 64.7%, which is higher than GPT-5.5 (58.6%) and nearly on par with Opus 4.7 (64.3%), but lower than Opus 4.8 (69.2%) and Fable (80.4%). In Terminal Bench 2.1, Grok 4.5 demonstrated 83.3% — almost identical to GPT-5.5 (83.4%) and slightly behind Fable (84.3%). However, on SWE Marathon, the model took first place with 29%, ahead of Opus 4.8 (26%) and Fable (24%). Leadership on Harvey's Legal Agent Benchmark was also claimed.
Token Economy: Where SpaceXAI Truly Wins
The strongest argument in favor of Grok 4.5 is not raw scores, but token usage efficiency. On SWE Bench Pro tasks, the model consumed an average of 15,954 output tokens per task. For Claude Opus 4.8, this figure was 67,020 tokens — 4.2 times more. With a lower price per token, this means a radical reduction in final cost for scenarios with a large number of iterations: bug fixing, generating edits, code review, and agent cycles.
SpaceXAI also claims a generation speed of about 80 tokens per second and a context window of 500,000 tokens. The model is positioned as a compromise between quality, speed, and cost — exactly what is needed for large-scale industrial deployments.
Connection with Cursor and Strategic Context
Notably, Grok 4.5 was trained in close collaboration with the AI service Cursor, which SpaceX recently agreed to acquire for $60 billion. The training used data from Cursor developer sessions, including debugging traces and real code edits, not just static repositories. This explains the model's strong focus on practical programming tasks.
To train Grok 4.5, the company utilized tens of thousands of Nvidia GB300 GPUs. The model is one of the first major ones following the merger of Elon Musk's AI division with SpaceX.
My expert opinion: SpaceXAI is betting not on being the best in all tests, but on being the most cost-effective in real-world usage scenarios. In an era where inference costs are becoming a critical factor for scaling AI, this approach may prove more successful than chasing abstract benchmark records. If Grok 4.5 truly delivers comparable quality at 4 times lower token costs, it could significantly reshape the AI model market for the enterprise sector.