MilikMilik

Grok 4.5 Joins the Opus-Class: A Cheaper Challenger to GPT-5.6

Grok 4.5 Joins the Opus-Class: A Cheaper Challenger to GPT-5.6
Interest|High-Quality Software

Grok 4.5 in One Line: Opus-Class Power, Price as the Weapon

Grok 4.5 is SpaceXAI’s new large AI model, jointly built with coding startup Cursor, positioned as an Opus-class system for coding and knowledge work that promises competitive benchmark performance with lower cost-per-token than top Anthropic and OpenAI models, aiming squarely at developers who care more about throughput and pricing than leaderboard glory.

This launch is not a lab demo; it is a pricing and positioning shot at the current frontier. SpaceXAI, formed after SpaceX absorbed xAI, has released Grok 4.5 as its first collaborative model with Cursor and its first major model launch since going public. It goes public on 9 July, the same day OpenAI opens GPT-5.6 Sol, Terra, and Luna to the public, putting Musk’s camp and OpenAI into their most direct model-versus-model fight so far. For developers, the key takeaway is blunt: if Grok 4.5 is “good enough,” its economics could matter more than small benchmark gaps.

Grok 4.5 Joins the Opus-Class: A Cheaper Challenger to GPT-5.6

Opus-Class Performance: Strong Benchmarks, Not the Uncontested Leader

SpaceXAI insists Grok 4.5 is “Opus-class,” roughly comparable to Anthropic’s top Opus family while being faster and more token-efficient. Benchmarks support part of that story: Grok 4.5 does not dominate, but it is now clearly in the top cluster of general-purpose and coding models. On Terminal-Bench 2.1, it scores 83.3%, beating Opus 4.8’s 78.9% and landing within a rounding error of GPT-5.5’s 83.4%. On DeepSWE 1.0 from Artificial Analysis, it scores 62.0%, ahead of Opus 4.8 at 55.8% and just behind GPT-5.5’s 64.3%.

The picture is less flattering on broader coding tests. Grok 4.5 trails Opus 4.8 on SWE-Bench Multilingual (78.0% vs 84.4%) and SWE-Bench Pro (64.7% vs 69.2%), and it lags Fable 5 across all reported categories, including Fable’s 80.3% on SWE-Bench Pro. Benchmark results released with the model describe it as competitive yet still slightly behind the strongest systems in some areas. The honest conclusion: Grok 4.5 is a serious top-tier option, but not a clean benchmark winner—its appeal has to come from cost and workflow fit.

Grok 4.5 Joins the Opus-Class: A Cheaper Challenger to GPT-5.6

Cursor Data and Monthly Flagships: A New Model-Building Strategy

What makes Grok 4.5 strategically interesting is how it was built, not just how it scores. It is the first model jointly trained by SpaceXAI and Cursor, using a mixture-of-experts architecture and trillions of tokens drawn from real Cursor user sessions, capturing how developers and coding agents work in actual codebases. That is a major shift from Cursor’s earlier Composer 2.5 coding specialist and signals a move toward models grounded in real developer telemetry rather than mostly synthetic or scraped data.

Grok 4.5 runs on xAI’s new V9 foundation with 1.5 trillion parameters, roughly three times the size of the v8-small architecture behind Grok 4.3. A developer at SpaceXAI has said this is the first flagship under the merged unit and that the plan is a from-scratch foundation model shipping every month through the end of 2026, with the next model already in training and set to include Cursor data from the very start of pre-training. In other words, Grok 4.5 is both a product and a proof of process: an AI built directly from live coding behavior.

Pricing and Token Efficiency: Where Grok 4.5 Tries to Beat GPT-5.6

If benchmarks are a draw, price is where SpaceXAI wants to win. Grok 4.5 is priced at USD 2 (approx. RM9) per million input tokens and USD 6 (approx. RM28) per million output tokens, with a faster variant at USD 4 (approx. RM18) input and USD 18 (approx. RM83) output. That undercuts Anthropic’s Opus 4.8 at USD 5 (approx. RM23) input and USD 25 (approx. RM115) output, and it sits near OpenAI’s GPT-5.6 Luna, which is priced at USD 1 (approx. RM5) per million input tokens and USD 6 (approx. RM28) per million output tokens. Another report notes Grok 4.5 at the same USD 2/6 levels against Opus 4.7, while OpenAI’s input pricing ranges from USD 1 to USD 5 (approx. RM5–RM23) depending on tier.

SpaceXAI claims Grok 4.5 delivers roughly twice the token efficiency of competing models, and on SWE-Bench Pro it says the model uses around 4.2 times fewer output tokens than Opus 4.8 while running at about 80 tokens per second. In a market where enterprise developers pay for billions of tokens, this matters more than a few extra benchmark points. The release lands in the same week OpenAI rolls out GPT-5.6 more broadly, intensifying a pricing and performance battle across coding workflows and turning cost-per-token into a primary competitive lever rather than a footnote.

What GPT-5.6 vs Grok Means for Developers Right Now

The timing alone says everything about the stakes. Grok 4.5 goes public on 9 July, the same day OpenAI opens GPT-5.6 Sol, Terra, and Luna to the public, after earlier export-related limits on GPT-5.6 had delayed a wider rollout. Musk pushed the launch based on strong positive feedback from a beta across SpaceX and Tesla, reviving his rivalry with OpenAI in the most direct form since he left its board; the competing launches this week put their strongest models head-to-head.

For developers, this rivalry is less about personalities and more about practical trade-offs. Grok 4.5 is aimed squarely at coding, agentic workflows, and high-value knowledge work in finance, legal, office tasks, research, and writing, promising to handle long-running, tool-heavy tasks across software engineering, data science, finance, and legal work. OpenAI, meanwhile, pairs GPT-5.6 with GPT-Live for voice, countering SpaceXAI’s Voice Agent Builder. In the short term, the smart move is not to crown a single winner but to treat Grok 4.5, GPT-5.6, and Anthropic’s models as a portfolio: route workloads by cost, latency, and task type. The clear conclusion: the era of one default “best model” is over—pricing and task fit will decide your stack.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!