xAI shipped Grok 4.5 on July 8, describing it as the company’s first model trained specifically for coding and agentic work, positioned as a workhorse capable of frontier-level performance at a fraction of the cost of rivals. Independent testing over the following two days has largely confirmed the cost claim and complicated almost everything else. Grok 4.5 is a genuinely strong, efficient model. It is not, by the measures independent evaluators actually use, the best model on the market, and in several head to head coding benchmarks it trails Anthropic’s and OpenAI’s current flagships by a meaningful margin.
That gap between framing and evaluation is the real story here, more than any single benchmark number.
What Independent Evaluators Found
Artificial Analysis, one of the more widely cited third-party evaluation groups, gives the clearest single snapshot. Its July 8 analysis puts Grok 4.5 at 54 on its Intelligence Index, a composite score spanning reasoning, knowledge, and task completion. That’s a 16-point jump over Grok 4.3 and represents genuine progress, but it lands the model fourth overall, behind Claude Fable 5 at 60, Claude Opus 4.8 at 56, and GPT-5.5 at 55. The gap to the top of that list is real, even if the gap to third place is close enough to call a rounding error.
The coding-specific benchmarks tell a similar story with sharper edges. On DeepSWE 1.1, which measures how reliably a model resolves real developer-submitted bugs, Grok 4.5 scored 53%, trailing GPT-5.5’s 67% and Claude Fable 5’s 70% by a wide margin. On SWE-Bench Pro, a broader software engineering test, Grok 4.5 posted 64.7% enough to edge out GPT-5.5’s 58.6%, but still well short of Opus 4.8’s 69.2% and Fable 5’s 80.4%. Terminal-Bench 2.1, which tests multi-step command-line reasoning, was the closest race of the group: Grok 4.5 scored 83.3%, essentially tied with GPT-5.5’s 83.4%, and only about a point behind Fable 5’s 84.3%.
Where Grok 4.5 Doesn’t Trail
None of this means Grok 4.5 is a weak model, and reading the benchmarks selectively in either direction misses what’s actually interesting about the release. Artificial Analysis found Grok 4.5 takes the single top spot on independent agentic tool-use testing, meaning it selects the right tool more consistently and recovers from errors better than any other model tested a genuinely distinctive result given how much modern coding and research work depends on chaining tool calls reliably. Separately, evaluation firm Snorkel ran Grok 4.5 against GPT-5.5 and Opus 4.8 on its GDPval+ dataset of expert-authored professional workplace tasks, and Grok 4.5 came out ahead on aggregate, posting a 29% pass rate against 22% for GPT-5.5 and 21% for Opus 4.8, with the largest gaps showing up in legal work, education, and healthcare tasks.
Cost is where the model’s positioning is least disputable. Artificial Analysis reports a cost of roughly $0.31 per Intelligence Index task, well below its immediate rivals, and xAI’s own figures claim Grok 4.5 uses roughly 4.2 times fewer output tokens than Opus 4.8 on SWE-Bench Pro tasks. Independent measurement broadly supports that direction, even if the exact multiplier varies by source. There is a real caveat buried in the same data, though: Artificial Analysis also found that while Grok 4.5’s accuracy on its Omniscience knowledge benchmark rose from 35% to 52% over the previous Grok generation, its hallucination rate more than doubled, from 25% to 54% a sign the model has become both more knowledgeable and more confidently wrong at the same time.
The taken-together picture is a model that is not the strongest generalist available, but is a genuinely strong, cheap, agentic option closer to “excellent value at the frontier’s edge” than “new best model,” whatever the launch messaging implied.
The Marketing Behind the Numbers
Context matters here, and xAI’s has been unusually turbulent. The company’s AI division, folded into a newly public SpaceX earlier this year, had lost all eleven of its original co-founders by the end of March, and Musk has publicly acknowledged the unit needed to be rebuilt “from the foundations up.” SpaceX’s roughly $75 billion IPO in June was followed within days by a $60 billion all-stock deal to acquire Cursor, the AI coding startup, explicitly aimed at catching up to Anthropic and OpenAI in coding tools after Grok’s earlier models struggled to compete. Grok 4.5 is the first major model release since that acquisition, and xAI has said it was trained in collaboration with Cursor using real developer session data rather than static code repositories alone.
That context helps explain some of the specific choices in how xAI presented the launch. The company’s own comparison charts benchmarked Grok 4.5 against GPT-5.5 rather than GPT-5.6, even though the latter launched within hours of Grok’s announcement a comparison that was accurate at the moment it was drawn up, but one that aged out of relevance almost immediately. xAI also published results for a narrower set of benchmarks than its rivals typically release; where GPT-5.5 has verified scores across more than a dozen published benchmarks spanning browsing, tool orchestration, and scientific reasoning, Grok 4.5’s official materials leaned heavily on the coding metrics that presented the model most favorably.
None of that is unusual for a product launch. It is, however, exactly why independent, third-party evaluation matters more than any vendor’s own comparison chart, and why a company’s framing of “our newest model, tested against the field” deserves scrutiny before it gets repeated as settled fact.
The Bottom Line
Grok 4.5 is a legitimate frontier-adjacent model, not a marketing mirage it leads on agentic tool use, undercuts its rivals sharply on cost, and even topped a well-regarded professional tasks benchmark against GPT-5.5 and Opus 4.8. What it isn’t, based on the independent evaluation that has emerged since launch, is the best model available by the measures that matter most for raw coding accuracy and general capability, where it trails Claude Fable 5, Claude Opus 4.8, and in several benchmarks, GPT-5.5 as well. The gap between those two realities isn’t a contradiction. It’s what happens when a company under real competitive pressure, rebuilding its AI division after a rocky year, ships a genuinely useful model and frames it as more than that. For anyone deciding whether to route work to Grok 4.5, the honest answer is that it’s an excellent choice for high volume, cost-sensitive agentic tasks, and a weaker one for raw coding accuracy where budget isn’t the constraint a more useful takeaway than any single benchmark chart, xAI’s or otherwise, can offer on its own.







Leave a Reply