Grok 4.5: A Coding-First Flagship Built to Win on Economics
Grok 4.5 is a new flagship AI coding model from SpaceXAI that targets developers and office workers with frontier-level coding benchmarks, a huge 500,000-token context window, and aggressive API pricing designed to lower the real cost of completing software and document workflows. SpaceXAI launched Grok 4.5 on July 8 as a frontier model for coding, agentic workflows, and knowledge work, calling it the company’s most capable model so far. Rather than chase the very top benchmark slot at any cost, SpaceXAI is making a clear bet: developers care more about cost per finished pull request or spreadsheet than raw IQ points.
The Grok 4.5 model supports text and image input, text output, function calling, structured outputs, web and X search, code execution, and configurable reasoning within a 500,000-token window. It launches directly where developers live: Grok Build, the API console, and Cursor on all plans. This is not framed as a chatbot refresh, but as a coding and agent engine meant to sit under IDEs, CLIs, and Office add-ins.

Pricing and Benchmarks: Frontier-Class Coding Without Frontier Prices?
The headline play is price. Grok 4.5 comes in at USD 2 (approx. RM9.20) per 1 million input tokens, USD 0.50 (approx. RM2.30) per 1 million cached input tokens, and USD 6 (approx. RM27.60) per 1 million output tokens. According to SpaceXAI, “Grok 4.5 is priced at $2 per million input tokens and $6 per million output tokens, well below rival flagships.” The company’s own comparison lists Anthropic’s Claude Opus 4.8 at USD 5 (approx. RM23) input and USD 25 (approx. RM115) output, and OpenAI’s GPT-5.6 Luna at USD 1 (approx. RM4.60) and USD 6 (approx. RM27.60).
On coding benchmarks, Grok 4.5 is in striking distance of the frontier. It scores 62.0% on DeepSWE 1.0, 53% on DeepSWE 1.1, 83.3% on Terminal Bench 2.1, and 64.7% on SWE Bench Pro. It also ranks #4 on GDPval-AA v2 with an Elo of 1543, behind only the latest Claude releases on real-world agentic knowledge work tasks. SpaceXAI argues that the model wins not only on raw price, but on cost per completed task; one analysis notes Grok 4.5 used 15,954 output tokens per SWE Bench Pro task versus 67,020 for Claude Opus 4.8 in the same setup, a huge gap if it holds in production.

Cursor Integration Turns Grok 4.5 Into a Default AI Coding Assistant
The strategic move is where Grok 4.5 shows up. SpaceXAI is not asking developers to switch chatbots; it is inserting the model straight into existing workflows. Grok 4.5 is live in Grok Build, in the SpaceXAI API console, and inside Cursor across desktop, web, iOS, CLI, and SDK, with usage included on individual and team plans. SpaceXAI recently agreed to buy Anysphere, the company behind Cursor, in a roughly USD 60 billion (approx. RM276 billion) all-stock deal, and Grok 4.5 is the first big product to reflect that alignment.
This tight Cursor integration matters more than any glossy launch demo. Cursor’s interaction data was used to train the Grok 4.5 model, feeding it signals on how engineers write, review, and debug code in practice. That gives Grok 4.5 a natural path to become the default AI coding assistant inside one of the most popular AI-first IDEs, a direct challenge to incumbent assistants tied to other leading models. For teams already living in Cursor, trying Grok 4.5 is now the path of least resistance.
From Coding Agents to Office Workflows: Who Benefits First?
Grok 4.5 is pitched first as an engine for AI coding agents and long-running tool use. It is SpaceXAI’s “first model trained specifically for coding and agents,” and can act as a coding agent in Grok Build through the CLI, terminal UI, scripts, bots, and Agent Client Protocol integrations. The model can handle languages like Rust and C/C++, and claims to build entire apps from a single prompt, aiming to reduce retries and prompt fiddling that inflate costs in rival systems.
But the same engine is being pushed into knowledge work. Grok 4.5 runs inside Office add-ins for Word, PowerPoint, and Excel, where it can create spreadsheet models, slide diagrams, and document drafts, with support for advanced multi-sheet spreadsheets and web-enabled workflows. That should translate to quicker responses for everyday users, while developers and businesses using the API could also benefit from lower operating costs if fewer retries and shorter outputs hold up outside benchmarks. In practice, the big winners will be teams willing to redesign workflows around agents rather than treat the model as a fancier autocomplete.
Limits, EU Lag, and the Real Test: Cost per Finished Task
Grok 4.5 is not without caveats. EU access is still pending; the model is not available in EU products or the API console yet, with availability expected in mid-July. Reasoning effort is configurable but cannot be disabled, with “high” set by default, which may concern teams who want precise control over costs. SpaceXAI also warns that higher-context pricing can apply above 200,000 tokens within the 500,000-token window, and server-side tools like web search, X search, and code execution can add extra fees.
The harder truth is that token prices alone never decide the winner. Grok 4.5 is cheaper than some Opus-class rivals on paper, but more expensive than older models and open-weight options. If it needs fewer tokens, fewer retries, and fewer tool calls to ship a pull request or finish a complex Excel model, it will earn its place in stacks; if not, the savings evaporate. For now, Grok 4.5 looks like a credible, developer-first alternative in the AI coding assistant race—and a clear signal that the next phase of AI competition will be about economics and integration, not only intelligence scores.






