Elon Musk agreed “for now” with an AI leaderboard chart that ranks his Grok models behind both Anthropic and OpenAI, according to reports circulating September 28, 2026. The chart, which spread widely on social media before Musk weighed in, places Anthropic’s Claude Opus 5.5 at the top of the field, with OpenAI’s GPT-6 Astra in second. Musk’s own models, Grok 4.6 and Grok 4.7, land three tiers lower, and the older Grok 4.5 sits further down still.

It’s a notable moment for a founder who has spent much of 2026 promising that Grok would “exceed every model” on the market. The admission, brief as it was, lands the same week that independent benchmarking firm Artificial Analysis published its own scoring of Grok 4.7, giving outside observers a second, numbers-based way to check whether Musk’s concession matches the data.

Musk Endorses a Chart That Puts Grok Behind Rivals

The story traces back to a ranking graphic shared online that sorted major AI labs into tiers based on model capability. Musk’s response, described by outlets including Longbridge and Hong Kong’s Futu News, was that the ranking was “accurate for now.” That phrasing matters. Musk didn’t dispute the placement, and he didn’t concede it as permanent either. He left the door open for Grok to climb, which is consistent with how xAI has talked about its release cadence all year: ship fast, iterate faster, and treat any single leaderboard snapshot as temporary.

The companies named in the coverage are Anthropic, OpenAI, Space Exploration Technologies Corp., and Alphabet Inc., though the report’s specifics on methodology, exact publication source, and full date weren’t independently confirmed across every outlet that picked up the story. Readers should treat the underlying chart as a snapshot of sentiment and benchmark data rather than an official, audited ranking from a single standards body.

What the Ranking Chart Actually Shows

According to the reporting, the tiers broke down like this: Claude Opus 5.5 sits alone at the top. GPT-6 Astra follows in second place. A third tier includes Fable 5.1 and GPT-6 Sol, one rank below the leaders. Grok 4.6 and Grok 4.7 land three tiers below the top-ranked model, and Grok 4.5 falls even further behind. That’s a wide spread for a market where xAI has spent heavily on compute and talent to close the gap with the two biggest labs.

It’s worth being precise about what this kind of chart measures versus what it doesn’t. A tiered leaderboard usually blends multiple benchmark categories, sometimes weighted by task type, sometimes by user preference data. It’s not the same thing as a single, apples-to-apples score. That distinction is exactly why Grok’s position looks different depending on which benchmark you check, a pattern that shows up clearly once you compare the tiered chart against Artificial Analysis’s own numbers.

Grok 4.7’s Numbers: Where the Gains Are Real

Musk’s “for now” comment doesn’t mean Grok stood still. Artificial Analysis, an independent benchmarking firm, scored Grok 4.7 at 1,657 Elo on its AA-Briefcase test, which evaluates models on realistic, long-horizon professional work tasks. That’s a jump of 111 points over Grok 4.6, and it places Grok 4.7 just behind Claude Opus 5 and Claude Fable 5.1 on that specific measure. On the firm’s broader Intelligence Index, Grok 4.7 scored 46, two points above Grok 4.6, with Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5 all ranking higher.

When evaluated inside its own native harness with Grok Build, Grok 4.7 ranked fourth overall, trailing Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. So depending on which test you pull up, Grok 4.7 is either “just behind” the leaders or a clear step down, and both descriptions are technically accurate. That gap between narratives is the real story here, more than any single tier placement on a shared chart.

Pricing Stayed Flat Even as Capability Moved

xAI kept Grok 4.7’s pricing at $2 per million input tokens and $6 per million output tokens, the same rate as Grok 4.6. That’s a deliberate call: rather than charging more for the performance bump, xAI held price steady while pushing the underlying model harder, a strategy that lines up with the broader price compression happening across the frontier AI market this year, something we covered when Grok 4.7 shipped at $2/$6 while trailing GPT-6 by 34 points.

How Claude Opus 5.5 Took the Top Spot

Anthropic’s climb to the top of the chart didn’t happen overnight. The company has spent 2026 pushing Claude’s coding and agentic capabilities hard, and Opus 5.5 built on that work with a full 1-million-token context window aimed squarely at professional developers, a detail we broke down when Claude Opus 5.5 shipped 1M-token context for coding tools. Anthropic also positioned the release on cost, not just capability, which we covered in our look at how Claude Opus 5.5 beat GPT-5.6 Sol while costing a third as much to run.

That combination, leading benchmark scores plus a lower price than rivals, is a harder position for competitors to attack than a pure capability lead would be. It’s one reason Anthropic’s placement at the top of this particular chart tracks with how the company has framed its own releases all year: capability first, cost-efficiency as the closer.

OpenAI’s GPT-6 Astra and the Fight for Second Place

GPT-6 Astra’s second-place slot on the chart comes after a year in which OpenAI pushed hard on both frontier capability and, separately, security posture for its model family. Astra scored well enough on exploit-focused testing that we reported it beating an internal predecessor 88% to 12.5% on exploit tests, and Anthropic reportedly responded by weighing a new release once GPT-6 Astra’s score moved past a 13% threshold on a benchmark the two labs have been trading blows on. GPT-6 Sol, a lower-cost sibling model, landed one tier below Astra on the same chart, alongside Fable 5.1, illustrating that OpenAI and Anthropic are now both fielding multi-tier model families rather than single flagship releases.

Why “For Now” Matters More Than It Sounds

Two words carry a lot of weight in a market where model rankings shift monthly. Musk has built Grok’s entire public narrative around speed of iteration, arguing that raw compute scale and rapid release cycles will eventually close any capability gap. Agreeing that a ranking is accurate “for now” fits that narrative without abandoning it. It’s an admission of the present standing, paired with an implicit claim about the trajectory.

That framing isn’t new for Musk. He’s made similar claims about Grok catching up to rivals before, including teasing a much larger model earlier this year, a story we covered when Musk teased a 3-trillion-parameter Grok and reports split on the specs. The pattern suggests Musk is comfortable losing a given month’s leaderboard round as long as he can point to a bigger release on the horizon.

Benchmark Volatility: Why Rankings Swing by Methodology

One detail underscores how shaky single-chart rankings can be as a source of truth. A separate benchmark page tracking dozens of models listed Grok 4.6 in 15th place overall with a supported score of 69.19, a very different picture from a three-tiers-below-the-leader placement. Neither number is wrong exactly, they’re measuring different things, weighted differently, against different task sets. That’s the core problem with treating any single leaderboard as gospel: methodology choices move rankings by double-digit places without any change in the underlying model.

For engineering teams choosing a model for production, that volatility is the actual takeaway. A model that looks mid-tier on a consumer-facing chart can lead on the specific task category your product depends on, whether that’s code generation, long-context retrieval, or agentic tool use. Checking the benchmark’s task composition matters more than checking its headline rank.

The Same Models, a Different Order

A separate benchmark tracker, Artificial Analysis’s dedicated Grok 4.7 model page, breaks the same release down by individual task category rather than a single blended score. Reading the per-task breakdown alongside the tiered chart is the fastest way to see why “three tiers behind” and “just behind the leaders” can both be defensible descriptions of the same model, depending on which slice of the data you’re looking at.

Historical Context: From Grok 4’s Early Lead to Today

Grok’s position relative to Claude and GPT has swung more than once. After Grok 4 launched, Artificial Analysis briefly ranked it as the most intelligent model among leading competitors, ahead of the era’s Claude and GPT releases. That lead didn’t hold. By mid-2026, reports tracking Grok’s download numbers found them falling as Anthropic’s Claude widened its lead in both capability and adoption. Grok 4.6 then closed some of that distance in August, matching OpenAI’s GPT-5.6 Sol Max on the Intelligence Index according to Artificial Analysis, trailing only Anthropic’s top models at the time.

Grok 4.7’s September release, and Musk’s own acknowledgment of where it currently sits, is really just the latest data point in that back-and-forth. The three companies have effectively been trading the top slot in different benchmark categories for over a year, and no single lab has held a clean, uncontested lead across every measure at once.

Competitive Comparison: Grok, Claude, and GPT-6 Side by Side

The table below pulls together the publicly reported figures on pricing and Artificial Analysis benchmark scores for the models named in the September 2026 ranking discussion.

ModelMakerReported TierAA-Briefcase EloPricing (per 1M tokens, in/out)
Claude Opus 5.5AnthropicTop tierHighest among named modelsSee Anthropic’s 2026 pricing update
GPT-6 AstraOpenAISecond tierAbove Grok 4.7Not fully disclosed in reports
Fable 5.1 / GPT-6 SolAnthropic / OpenAIThird tierAbove Grok 4.7Varies by product line
Grok 4.7xAIThree tiers below top1,657 (up 111 from Grok 4.6)$2.00 / $6.00
Grok 4.6xAIThree tiers below top1,546$2.00 / $6.00
Grok 4.5xAIBelow Grok 4.6/4.7Not disclosed in reportsNot disclosed in reports

Note that the “reported tier” column reflects the leaderboard chart Musk responded to, while the Elo figures come from Artificial Analysis’s separate, independently run AA-Briefcase benchmark. The two data sources agree directionally, Grok trails the Anthropic and OpenAI flagships, but they disagree on how close that gap actually is.

Where Each Lab Has Publicly Focused Its 2026 Roadmap

LabFlagship Model2026 Public FocusLatest Price Move
AnthropicClaude Opus 5.5Coding, 1M-token context, cost-per-token efficiencyCut price roughly 20%, cache cost roughly 60%
OpenAIGPT-6 AstraSecurity testing, exploit resistance, agentic reliabilityLaunched GPT-6 Sol and Luna near 50% below prior pricing
xAIGrok 4.7Compute scale, rapid iteration, agentic knowledge workHeld pricing flat at $2/$6 versus Grok 4.6

See Anthropic’s own newsroom and this 2026 rundown of xAI’s Grok releases, the SpaceX merger talk, and Colossus buildout for more background on how each lab has framed its roadmap this year.

Market and Business Impact for xAI

xAI has tied its business case tightly to Musk’s other ventures, particularly the compute and data advantages that come from proximity to Musk’s broader corporate holdings. Some market coverage, including from Benzinga, has referred to Musk’s AI venture under a SpaceX corporate designation in its financial reporting, though that specific corporate and ticker framing hasn’t been independently verified through an official company or exchange filing in the sources reviewed for this story. Readers should treat any specific stock ticker tied to xAI or SpaceX as unconfirmed until a primary source, such as an SEC filing or an official company statement, verifies it.

What is clear is that Grok’s download numbers took a hit earlier in 2026 as Claude’s adoption widened, based on figures reported in the spring. A public leaderboard admission from Musk himself, even a hedged one, is the kind of headline that can shape enterprise buying conversations in the short term, regardless of what the underlying benchmark data shows on a task-by-task basis.

Pricing Pressure Across the Frontier Model Market

The ranking debate is unfolding against a backdrop of aggressive price cuts across every major lab. OpenAI launched GPT-6 Sol and GPT-6 Luna at roughly half off prior pricing, a move we detailed in our coverage of how GPT-6 Sol and Luna launched at 50% off to undercut Claude. Grok 4.7 held its price steady rather than cutting further, which is itself a signal: xAI appears to be competing on capability-per-dollar rather than joining a race to the bottom on headline pricing.

That pricing backdrop matters for how buyers should read the leaderboard story. A model that ranks lower on a tiered capability chart but holds pricing steady while rivals cut theirs is making a different kind of bet: that developers will pay for consistency and predictable performance rather than chasing whichever model tops this week’s chart.

What This Means for Enterprise AI Buyers

Procurement teams evaluating frontier models in Q4 2026 should treat this leaderboard story as a reason to run their own task-specific benchmarks rather than a reason to switch vendors outright. The AA-Briefcase data shows Grok 4.7 landing close behind Claude Opus 5 and Claude Fable 5.1 on realistic professional work tasks, which is a meaningfully different signal than “three tiers behind” on a general capability chart. Teams building coding-heavy or agentic workflows should weight benchmarks that test those specific capabilities, not aggregate leaderboard placement.

It’s also worth watching how xAI responds over the next release cycle. If Musk’s “for now” framing holds, expect another Grok release within a few months aimed squarely at closing the AA-Briefcase gap with Anthropic’s top models, likely without a price increase given how xAI has priced Grok 4.6 and Grok 4.7 identically.

Predictions: Where the Leaderboard Goes Next

  • Grok will likely narrow the AA-Briefcase Elo gap with a follow-up release before the end of 2026, given the 111-point jump already logged between Grok 4.6 and Grok 4.7.
  • Anthropic and OpenAI will keep trading the top two leaderboard slots month to month rather than one lab establishing a durable lead, continuing the pattern seen since Claude Opus 5.5 and GPT-6 Astra shipped.
  • Pricing will stay the more decisive competitive lever short-term, with GPT-6 Sol and Luna’s roughly 50% price cuts putting pressure on xAI and Anthropic to either match or double down on a premium-quality position.
  • Expect more “single chart” leaderboard controversies as long as labs keep releasing on overlapping schedules, since any snapshot chart will look outdated within weeks given the current release cadence.
  • Enterprise buyers will increasingly demand task-specific benchmark disclosures (coding, agentic tool use, long-context retrieval) rather than accepting aggregate leaderboard tiers as a purchasing signal.

The Bigger Picture: A Three-Way Race With No Clean Winner

Strip away the headline and what’s left is a three-way race that’s been tightening and loosening in cycles for more than a year. Anthropic currently holds the top chart position with Claude Opus 5.5. OpenAI’s GPT-6 Astra sits close behind. Grok 4.7 trails both on the tiered chart Musk responded to, but closes much of that distance on Artificial Analysis’s own benchmark, particularly on realistic professional work tasks. None of that is likely to be settled by the time the next model ships from any of the three labs.

Musk’s two-word concession is memorable mostly because it’s rare. Lab founders don’t often agree publicly with a chart that ranks their own product behind competitors, hedge or not. That alone is probably why the story spread as fast as it did, more than the actual size of the capability gap it describes.

Frequently Asked Questions

Did Elon Musk say Grok is worse than Claude and GPT-6?
Reports describe Musk agreeing that a leaderboard chart placing Grok below Anthropic’s and OpenAI’s top models was “accurate for now.” That’s the confirmed wording; broader paraphrases suggesting he called Grok worse outright go beyond what’s been verified.

What ranks first on the chart Musk responded to?
Anthropic’s Claude Opus 5.5 was reported at the top, with OpenAI’s GPT-6 Astra second, according to the coverage reviewed for this story.

How far behind is Grok 4.7 really?
It depends on the benchmark. On the tiered chart, Grok 4.6 and Grok 4.7 sit three tiers below the top model. On Artificial Analysis’s AA-Briefcase test, Grok 4.7 scored 1,657 Elo, just behind Claude Opus 5 and Claude Fable 5.1.

Did Grok 4.7’s price change from Grok 4.6?
No. Both are priced at $2 per million input tokens and $6 per million output tokens, according to Artificial Analysis’s reporting.

Is the ranking chart an official industry standard?
No. It’s a chart that circulated publicly and drew a response from Musk. Independent benchmarks like Artificial Analysis’s Intelligence Index and AA-Briefcase are separate, methodology-disclosed measures that sometimes tell a different story than a single tiered chart.

Why do different benchmarks rank Grok so differently?
Benchmarks weight task categories differently. A leaderboard skewed toward general chat quality can rank a model very differently than one focused on coding, agentic tool use, or long-context professional tasks, which is why Grok 4.6 appeared in 15th place on one benchmark page while ranking closer to the top on another.

What does this mean for companies choosing an AI model?
Run task-specific evaluations rather than relying on a single leaderboard snapshot. A model that trails on aggregate capability charts can still lead on the specific workload, like coding or long-document retrieval, that matters for a given product.

Is xAI part of SpaceX?
Some market coverage has referenced Musk’s AI venture under a SpaceX corporate designation in financial reporting, but that specific corporate structure and any associated stock ticker have not been independently verified through an official company statement or exchange filing in the sources reviewed here.