SpaceXAI has officially introduced Grok 4.5 — a new language model focused on programming, agentic tasks, and intelligent work. The company positions it as its most powerful system to date. The model is already available via API, in Grok Build, and in the Cursor development environment across all pricing plans. Users in the European Union cannot access it yet — the launch there is scheduled for mid-July.
The key advantage of Grok 4.5, according to the developers, is the combination of performance with an aggressive pricing policy. The model costs $2 per 1 million input tokens and $6 per 1 million output tokens. For comparison: Anthropic's Claude Opus 4.8 costs $5 and $25 respectively, Claude Fable 5 costs $10 and $50, and OpenAI's flagship GPT-5.6 Sol costs $5 and $30. Within the same line, GPT-5.6 Luna offers an even lower price — $1 for input and $6 for output, but is in limited preview mode. This makes Grok 4.5 one of the most affordable solutions in its class.
Elon Musk described the new model as "Opus-class, but significantly faster, cheaper, and more token-efficient." According to him, internal SpaceXAI tests show that Grok 4.5 is "roughly comparable to Opus 4.7, but much faster."
Benchmark Results: Mixed but Competitive
The benchmark data published by SpaceXAI paints a mixed picture. On DeepSWE 1.0, the model scored 62%, trailing Fable (66.1%) and GPT-5.5 (64.31%), but surpassing Claude Opus 4.8 (55.75%) and Opus 4.7 (40.12%). On the more complex DeepSWE 1.1, the result was lower — 53% compared to 59% for Opus 4.8 and 70% for Fable. On SWE Bench Pro, Grok 4.5 scored 64.7%, higher than GPT-5.5 (58.6%) and on par with Opus 4.7 (64.3%), but lower than Opus 4.8 (69.2%) and Fable (80.4%).
However, on some tests, the model performed stronger. On Terminal Bench 2.1, it scored 83.3%, nearly matching GPT-5.5 (83.4%) and only slightly behind Fable (84.3%). On SWE Marathon, Grok 4.5 took first place with 29%, ahead of Opus 4.8 (26%) and Fable (24%). The company also claimed leadership on Harvey's Legal Agent Benchmark, a test for legal agentic tasks.
Economics — The Main Trump Card
The strongest argument in favor of Grok 4.5 lies not in peak performance, but in economics. SpaceXAI provides data showing that on SWE Bench Pro tasks, the model uses an average of 15,954 output tokens per task, while Claude Opus 4.8 uses 67,020 tokens — 4.2 times more. Combined with a significantly lower price per token, this makes Grok 4.5 extremely cost-effective for scenarios with a high number of iterations: bug fixing, generating edits, code review, and agentic loops.
Additionally, a generation speed of about 80 tokens per second and a context window of 500,000 tokens are claimed. SpaceXAI positions the model as the optimal compromise between quality, speed, and cost.
Connection with Cursor and Strategic Context
It is important to note that Grok 4.5 was trained jointly with the AI service Cursor, which SpaceX recently agreed to acquire for $60 billion. The deal will be conducted with Class A shares and is expected to close in the third quarter of 2026. The model was trained using data from real Cursor developer sessions, including debugging traces and code edits, giving it a unique advantage in understanding practical programming tasks. Tens of thousands of Nvidia GB300 GPUs were used for training.
This release is one of the first major initiatives following the merger of Elon Musk's AI division with SpaceX. Grok 4.5 does not attempt to be the undisputed leader in all tests, but offers a pragmatic approach: sufficiently high quality at a price that could reshape the economics of enterprise applications. In a world where inference costs are becoming the main barrier to scaling AI, this approach looks not just reasonable, but potentially dominant.