A five-prompt test that Tom’s Guide published on September 28, 2026 is doing something that dozens of benchmark leaderboards haven’t managed all year: it’s making people question whether ChatGPT is still the default answer to “which AI assistant should I use.” AI editor Amanda Caswell ran Meta Muse against ChatGPT-6 on five ordinary tasks, no coding riddles, no PhD-level math, just the stuff people actually ask an assistant to do. Muse won three of the five rounds. The headline called it what it was: “OpenAI just lost.”

That single test is now rippling through a market that has spent the back half of 2026 watching Meta’s consumer AI push gain real traction. Meta Muse’s download curve has already been outpacing ChatGPT’s early trajectory, and this head-to-head gives that momentum a concrete, human-readable data point instead of just an app-store chart. The test wasn’t run in a lab with synthetic prompts designed to flatter one model. It was five everyday requests, and the split decision tells you more about where the assistant war is actually headed than another leaderboard score would.

What Tom’s Guide actually tested

Caswell’s methodology was deliberately mundane. Five prompts, covering writing, planning, problem-solving, explaining, and creative reasoning, the categories most people actually lean on an AI assistant for during a normal week. No trick questions, no adversarial jailbreak attempts, no obscure math olympiad problems. Just the kind of requests that show up in a real inbox or a real group chat.

Three specific prompts from the test are worth walking through individually, because the pattern in the results is more interesting than the final score.

The neighbor’s trash cans prompt

The first head-to-head asked each assistant to draft a friendly text message about a neighbor’s trash cans left out too long, a small, socially delicate situation that requires reading a room rather than solving an equation. Meta Muse won this round. The result lines up with what Meta has been building toward since it first launched Muse: an assistant tuned for tone, relationships, and the soft skills that a pure reasoning engine tends to flatten.

The NYC family itinerary prompt

The second test flipped the script. The prompt asked for a Saturday itinerary in New York City for a family with three children, a $200 budget, and a fixed 10 a.m. to 7 p.m. window. ChatGPT won this round decisively. Constraint-heavy planning, juggling a dollar figure, a headcount, and a time block, is exactly the kind of structured task where a model built around step-by-step reasoning tends to pull ahead.

The pre-guest cleaning plan prompt

The third comparable prompt asked for a 45-minute cleaning plan before guests arrive. ChatGPT won again. Same pattern as the itinerary test: a fixed time budget and a sequencing problem, the kind of task that rewards a model optimized for logistics over one optimized for empathy.

Put together, Caswell’s own framing of the results was that ChatGPT held the edge on careful planning and tasks with hard practical constraints, while Meta Muse pulled ahead whenever understanding people, not numbers, was the actual job. That’s not a knockout blow for either company. It’s a map of where each assistant’s strengths sit, and it’s the kind of nuance that a bare “3-2” scoreline doesn’t capture on its own.

Why a 3-2 score is doing so much damage to OpenAI’s narrative

OpenAI has spent three years building a reputation as the default, unbeatable option for conversational AI. A narrow, split-decision loss to a rival’s assistant, on ordinary tasks rather than niche benchmarks, chips at exactly the perception that has let ChatGPT keep charging premium subscription prices while fending off Google’s Gemini, Anthropic’s Claude, and now Meta’s Muse. It doesn’t matter that the test only covered five prompts. What matters is that it’s the kind of test millions of ordinary users can run themselves in ten minutes, and now some of them will, expecting to see ChatGPT dominate the way it always has.

It’s worth being precise about what is and isn’t confirmed here. The headline’s framing, that “OpenAI just lost,” is Tom’s Guide’s editorial characterization of a five-prompt test, not evidence of a broader competitive collapse. ChatGPT-6 is the name used in the article, but there’s no confirmed official OpenAI announcement, release date, spec sheet, or price point attached to it in the public record right now. The same caution applies to Meta Muse: beyond the name used in the piece, there’s no confirmed product spec sheet for whatever build was tested. What is confirmed is the test itself, the categories, the specific prompts, and the 3-2 result. That’s still a meaningful signal, just not a verdict on either company’s entire roadmap.

Meta Muse’s momentum didn’t start with this test

This result lands on top of a Muse growth story that’s already been building for months. Download tracking has shown Muse matching or beating ChatGPT’s early adoption curve, and Meta has kept shipping features at a pace that suggests it’s treating this as a genuine platform fight, not a side project. Recent updates have added video avatars, email access, and Mac control to Muse, pushing it well past the simple chatbot category and into the same “AI that manages your digital life” lane OpenAI has been trying to occupy with ChatGPT’s own agentic features.

That expansion hasn’t gone unchallenged. Amazon has already blocked Meta’s Muse agent over concerns tied to how it operates on Amazon’s platforms, a sign that as these assistants gain more control over accounts, calendars, and communication tools, the platform-level pushback is becoming as important a storyline as the head-to-head benchmark results. Meta’s willingness to keep pushing Muse’s permissions further, even while facing friction from other tech giants, tells you this is a long-term bet, not a one-off product launch.

The AI assistant market’s competitive landscape, mapped

Meta Muse and ChatGPT aren’t fighting in an empty room. Anthropic’s Claude and Google’s Gemini are both competing for the same consumer attention, and both have had their own head-to-head moments against ChatGPT this year. Claude Opus 5.5 has posted its own benchmark wins against OpenAI’s models at a fraction of the cost, and pricing has become as much a battleground as raw capability. OpenAI’s own GPT-6 Sol and Luna models launched at steep discounts specifically to undercut Claude, which shows the company is well aware that price, not just benchmark scores, is now a lever competitors can pull.

The table below lays out how the major consumer-facing assistants stack up on the traits that actually drove Tom’s Guide’s test results, tone and empathy versus structured planning, since that’s the axis this particular comparison turned on.

AssistantParent companyReported strength in this test cycleReported weaker areaNotable recent move
ChatGPT-6OpenAIConstraint-heavy planning, budgets, schedulesReading social/emotional nuanceNamed in Tom’s Guide’s Sept. 27-28 test
Meta MuseMetaTone, social nuance, relationship-aware repliesNot directly tested on hard logistics winsAdded video avatars, email, Mac control
Claude Opus 5.5AnthropicBenchmark wins vs. GPT-5.6 Sol at lower costNot part of this specific 5-prompt testPositioned on cost-efficiency vs. OpenAI
GeminiGoogleHigh traffic and visit share vs. ChatGPTNot part of this specific 5-prompt testCrossed 2 billion visits milestone this year

None of this means Meta Muse has overtaken ChatGPT as the market leader. ChatGPT still carries the largest installed base and the deepest enterprise integration story of any assistant on this list. But the gap between “market leader” and “best at everything” is exactly what this test exposed, and that gap is where a challenger like Meta gets its opening.

Historical context: how we got to a 3-2 split decision

Meta’s consumer AI assistant didn’t arrive as a finished product. Muse has shipped features in waves over the past year, each one closing a specific gap with ChatGPT rather than trying to leapfrog it in one release. That iterative approach mirrors how ChatGPT itself grew, starting as a text-only chatbot and gradually picking up voice, browsing, memory, and now the kind of agentic account access that both companies are racing to ship safely.

OpenAI, for its part, has been the company setting the pace for most of the last three years, forcing Google, Anthropic, and Meta to react to its release cadence rather than the other way around. The GPT-6 family’s aggressive pricing moves this year are a direct product of that pressure finally reversing, at least partially, as rivals started matching or beating OpenAI on specific tasks rather than chasing it on every front simultaneously. A five-prompt test where a challenger wins three rounds is a small data point in that longer arc, but it’s the kind of small data point that compounds if it keeps repeating.

What this means for everyday users choosing an assistant

The practical takeaway from Caswell’s test isn’t “switch to Muse” or “stick with ChatGPT.” It’s that the right assistant now depends on what you’re actually asking it to do, which is a genuinely new situation for a market that spent 2023 through 2025 treating ChatGPT as the default answer to every AI question. If your daily use case leans toward drafting messages, handling social situations, or anything where reading the room matters more than hitting a number, Muse has a real case. If you’re mapping out a schedule with a hard budget or a fixed time window, ChatGPT still has the edge according to this test.

That’s a harder pitch for either company to market than “we’re the best AI,” but it’s a more honest one, and it’s likely to become the standard framing as more outlets run their own comparable tests.

Market and stock impact: why this matters beyond app reviews

Consumer AI assistant performance has become a genuine market-moving signal this year, not just a tech-press curiosity. Muse’s download numbers have already been cited as a factor in investor sentiment around Meta’s AI strategy, and a widely shared head-to-head win against ChatGPT-6 adds a qualitative data point to that download-count story. OpenAI, as a private company, doesn’t have a stock price to move on a single article, but its enterprise customers, investors in its funding rounds, and the analysts tracking its next raise all read coverage like this as a proxy for product momentum.

Meta, by contrast, is a public company, and every data point that supports a “Muse is actually competitive” narrative feeds directly into how Wall Street prices Meta’s broader AI bet. A viral, easily digestible comparison test that a competitor’s own AI editor ran and published is the kind of coverage that’s hard to buy with an ad budget. It reads as independent, which is exactly why it’s landing so hard. Yahoo Finance’s technology desk picked up the same comparison within a day, which is typically a sign a story has crossed from tech-press curiosity into a wider financial-news audience.

The limits of a five-prompt test

It’s worth being blunt about what this test can’t tell you. Five prompts is not a statistically rigorous benchmark, it’s a snapshot. A different set of five prompts, run on a different day, against different model versions, could easily produce a different score. Caswell’s own test design leaned into everyday tasks specifically because that’s what most readers care about, but that choice also means the test says very little about coding ability, technical reasoning, multilingual performance, or the kind of enterprise workloads that make up a huge share of both companies’ actual revenue.

The comparison also doesn’t establish anything about safety, privacy handling, or how each assistant behaves with sensitive account access, categories that matter enormously given how much account and email access both companies have been adding to their assistants recently. A win on “draft a friendly text about trash cans” says nothing about how either model handles a request touching financial data or health information.

Competitive comparison: where each company goes from here

OpenAI’s playbook after a result like this tends to be predictable: ship an update, publicize a benchmark win somewhere else, and let the news cycle move on. That’s worked for the company through several rounds of “ChatGPT loses to X” coverage over the past two years. The GPT-6 Sol and Luna pricing moves already show OpenAI is willing to compete on cost when it can’t win on every capability axis, and expect a similar response here, likely an update to whichever model version was tested, framed as closing the social-reasoning gap Caswell identified.

Meta’s task is different. It needs to convert a viral comparison win into sustained usage, not just a download spike, and Meta’s own AI product hub shows the company treating Muse as one piece of a much broader consumer AI push rather than a standalone experiment. The company has already shown it’s willing to push Muse’s feature set aggressively, adding capabilities like video avatars and account access faster than some partners are comfortable with, given the Amazon pushback. The bigger question for Meta is whether it can keep winning on “understanding people” while also closing the gap on structured planning tasks, because right now this test suggests it’s a genuinely split field rather than a clean win for either side.

Five predictions for how this plays out

1. Expect more head-to-head consumer tests, not fewer. Once one outlet’s five-prompt comparison goes viral, competing publications tend to run their own versions within weeks, and each one will carry its own scoreline.

2. OpenAI will lean harder into agentic and account-access features to differentiate on utility rather than tone. If Muse keeps winning on empathy, ChatGPT’s counter is likely to be depth of integration, not a personality overhaul.

3. Meta will keep shipping Muse features at the current pace, even at the cost of more platform friction like the Amazon block. A growth story built on momentum doesn’t slow down to smooth over one partner dispute.

4. Pricing will stay a bigger lever than raw capability scores for at least the next two quarters. The GPT-6 Sol and Luna discount strategy suggests OpenAI already sees cost as the more defensible edge when benchmark wins get split.

5. Expect both companies to publish their own benchmark results within the next month specifically addressing the categories Caswell’s test covered. Neither OpenAI nor Meta tends to let an unanswered “we lost” headline sit for long.

A quick reference: the five prompts and their outcomes

For readers who want the test results at a glance, here’s how the confirmed portions of the comparison broke down.

Prompt categorySpecific taskWinnerWhy it likely won
Social/writingFriendly text about neighbor’s trash cansMeta MuseStronger read on social tone and nuance
PlanningNYC family itinerary, $200 budget, 10 a.m.–7 p.m.ChatGPTStructured, constraint-heavy logistics
Problem-solving45-minute pre-guest cleaning planChatGPTFixed time budget and sequencing
ExplainingNot individually detailed in available reportingNot specified—
Creative reasoningNot individually detailed in available reportingNot specified—

The overall scoreline across all five prompts came out 3-2 in Meta Muse’s favor, according to Tom’s Guide’s published results.

What OpenAI and Meta haven’t confirmed

Neither company has issued an official statement responding directly to this specific test as of this writing. That’s not unusual, individual outlet comparisons rarely get a formal corporate response unless they escalate into a broader narrative. What’s also unconfirmed is any official specification sheet, pricing structure, or launch timeline for ChatGPT-6 as a distinct, named release separate from OpenAI’s existing GPT-6 family. The same goes for whatever specific build of Meta Muse was used in the test. Readers should treat the model names as reported by Tom’s Guide rather than as confirmed, dated product launches with their own spec sheets.

Why this test resonates when a hundred benchmarks didn’t

AI benchmark culture has produced an overwhelming number of leaderboards this year, MMLU variants, coding evals, reasoning suites, and outlets like The Verge and TechCrunch have spent much of 2026 covering that scoring race, but most of those numbers mean very little to someone deciding which assistant to open on their phone. A five-prompt test about trash cans and a family day trip cuts through that noise precisely because it’s legible. Anyone can read it, nod along, and mentally file away “Meta’s AI is better at reading people, OpenAI’s is better at planning.” That kind of simple, memorable takeaway travels further on social media and in casual conversation than any percentage-point improvement on a technical benchmark ever will, and it’s exactly why a single Tom’s Guide article is generating more competitive pressure than most of the industry’s formal evaluation reports.

Frequently asked questions

What did Tom’s Guide actually test?
AI editor Amanda Caswell ran Meta Muse against ChatGPT-6 on five everyday prompts covering writing, planning, problem-solving, explaining, and creative reasoning, and published the results on September 28, 2026.

Who won the test?
Meta Muse won three of the five prompts, and ChatGPT won two, for a final 3-2 result in Muse’s favor.

Which specific prompts did ChatGPT win?
ChatGPT won the New York City family itinerary prompt (a $200 budget within a 10 a.m. to 7 p.m. window) and the 45-minute pre-guest cleaning plan prompt, both tasks with hard time or budget constraints.

Which prompt did Meta Muse win?
Muse won the prompt asking for a friendly text message about a neighbor’s trash cans left out too long, a task that rewarded social and emotional nuance over logistics.

Is ChatGPT-6 an officially released OpenAI product?
The Tom’s Guide article names it as ChatGPT-6, but there is no confirmed official OpenAI launch announcement, spec sheet, or price for a product under that exact name at the time of this test.

Does this test mean Meta Muse is now better than ChatGPT overall?
No. The test covered five everyday prompts and doesn’t address coding, technical reasoning, enterprise workloads, safety handling, or account-access security, all areas where the picture could look very different.

How does this fit with Meta Muse’s download numbers?
Muse has already been posting download growth that outpaces ChatGPT’s early trajectory, and this test adds a qualitative result to that quantitative momentum.

Will OpenAI or Meta respond officially to this specific test?
Neither company had issued a direct public response as of this writing. Individual outlet comparisons of this kind don’t always get a formal corporate reply, though both companies have a track record of shipping competitive updates shortly after high-visibility comparisons.