A new benchmark published Thursday puts a number on a problem voice AI companies have mostly talked around: how often their text-to-speech systems botch the name of a drug. Synthio Labs released the results September 17, 2026, from San Francisco, under the name DOSE, short for Drug-name Oral Synthesis Evaluation. The company calls it the first public benchmark built specifically to test how accurately voice AI models pronounce medication names, and the numbers it produced are rough enough to matter far beyond a lab report.

DOSE ran nine text-to-speech models against 274 drug names, including 146 that were newly approved. Each name was read aloud inside a real clinical sentence rather than in isolation, which matters, because a model that nails a drug name on its own can still stumble when that name is buried in a stream of medical language. Every model tested saw its pass rate drop on the newly approved names, in some cases sharply. The story was picked up quickly by outlets including Yahoo Finance Singapore, extremetech.com, PYMNTS.com, Inshorts, finance.biggo.com, and NDTV Profit, all pointing to the same headline figure: leading voice AI models mispronounce up to one in three newly approved drug names.

What Synthio Labs’ DOSE Benchmark Actually Tested

DOSE is not a generic speech-quality test. It is narrowly built around one failure mode: whether a text-to-speech model can correctly say a drug’s name out loud. According to the release, Synthio Labs scored each of the nine models against a set of 274 drug names, split between established medications and 146 newly approved names. The newly approved group is the interesting part. Established drug names have had years, sometimes decades, to work their way into training data, closed captioning, pharmacy databases, and everyday speech. Newly approved names have not had that runway, and DOSE was designed specifically to find out what happens when a voice AI model meets one for the first time.

Each test sentence embedded the drug name in clinical context, the kind of phrasing a pharmacist, telehealth bot, or medication-reminder app might actually generate. That design choice matters for anyone trying to judge how the benchmark translates to the real world. A model reading “amoxicillin” as a single flashcard word is a different challenge than a model reading “the patient was prescribed amoxicillin twice daily” inside a longer, more natural utterance. Synthio Labs built DOSE around the second version, which is the version that actually ships in production voice assistants.

Why “Drug-Name Oral Synthesis Evaluation” Matters Right Now

Voice AI has quietly moved into places where a mispronounced word carries real weight. Pharmacy chains use voice bots for refill confirmations. Telehealth platforms use text-to-speech to read back prescribing information. Medication-reminder apps read drug names aloud to patients, some of whom are elderly, visually impaired, or managing a dozen prescriptions at once. None of that is new by itself. What is new is the pace at which drug names keep entering the pipeline: dozens of new medications win approval every year, each with a name engineered by branding teams to sound distinct from existing drugs, which often makes them harder, not easier, for a language model to guess phonetically from spelling alone.

That is the gap DOSE exposes. A voice AI model can be fluent, natural-sounding, and broadly accurate on the vocabulary it was trained on, and still fail the moment it hits a name that did not exist when its training data was assembled. Synthio Labs’ benchmark gives that gap a number for the first time in a public, comparative format, rather than leaving it as anecdotal complaints from pharmacists or scattered app-store reviews.

The Headline Number: Up to One in Three New Drug Names Mispronounced

The framing that spread across coverage of the release, that leading voice AI models mispronounce up to one in three newly approved drug names, comes from Synthio Labs’ own announcement. It is a fair summary of part of the data, but it is worth pulling apart, because the two named results in the release do not point to identical severity.

ElevenLabs’ eleven_v3 model passed 93.0% of established drug names but only 67.1% of newly approved ones, a drop of just under 26 percentage points. Turn that pass rate into a failure rate and it lands close to the “one in three” framing: roughly 33 out of every 100 new drug names tripped the model up. Google’s Gemini TTS told a worse story. It passed 89.1% of established names, then dropped to 61.6% on newly approved ones, a fall of nearly 28 percentage points. That works out to a failure rate closer to 38%, or nearly two in five, which is meaningfully worse than the “one in three” headline suggests. The release’s own framing, in other words, describes ElevenLabs’ result more precisely than it describes Google’s.

Across all nine models and the full 274-name set, general-purpose systems passed between 63.1% and 80.3% of drug names overall, according to the release. That range covers established and newly approved names combined, which is why the isolated newly-approved numbers for ElevenLabs and Gemini look so much worse in comparison. The gap between “overall” and “newly approved only” is the entire point of the benchmark.

ElevenLabs’ eleven_v3: A 26-Point Drop on New Drug Names

ElevenLabs has built its reputation on natural-sounding synthetic speech, and its eleven_v3 model still posted the strongest established-name score named in the release, at 93.0%. That number alone would read as a win. Paired with its 67.1% score on newly approved names, it instead reads as a case study in how quickly accuracy can erode once a model leaves familiar vocabulary. A 26-point swing between two subsets of the same benchmark is not a rounding error. It is the difference between a model that a clinic could plausibly trust with routine medication scripts and one that needs a human double-check on anything recently approved.

ElevenLabs has not issued a public response to the DOSE results as of this writing, and Synthio Labs’ release does not include a rebuttal or comment from any of the tested vendors, even as the broader AI and machine learning industry keeps shipping new voice and language models at a rapid clip. That absence of vendor pushback, at least so far, leaves the numbers standing largely uncontested days after publication.

Google’s Gemini TTS Falls Further, to 61.6%

Gemini TTS is the model with the widest gap disclosed in the release. Its 89.1% established-name score put it in the same neighborhood as ElevenLabs, but its newly-approved score of 61.6% is the lowest of the two named results, and by a meaningful margin. This is the score that most undercuts the “one in three” headline, because a 61.6% pass rate does not describe one-in-three failing, it describes closer to two-in-five failing.

For a company that has spent much of 2026 pushing Gemini deeper into consumer workflows, including recent moves like Gemini’s transcription rollout inside Gmail, a public benchmark showing its text-to-speech system struggling with newly approved drug names is an awkward data point to sit alongside that expansion. Voice features tend to get bundled together in a platform’s marketing, even when the underlying models serve very different tasks, and a headline about medication mispronunciation does not distinguish neatly between Gemini’s transcription work and its speech synthesis work in the eyes of a general audience.

Microsoft Azure Struggles Most With Generic Drug Names

The release singles out Microsoft Azure for a different, arguably more concerning result: the company’s text-to-speech system passed fewer than half of generic drug names, the kind known formally as International Nonproprietary Names, or INNs under the World Health Organization’s naming program. Generic names are not obscure or newly minted. They are the base chemical names that show up on nearly every prescription label, insurance form, and pharmacy printout in circulation. A sub-50% pass rate on that category is a different kind of problem than struggling with brand-new drugs, because it suggests the gap is not just about novelty. It is about a broader weakness in handling pharmaceutical vocabulary generally, on a platform, Azure AI Speech, that is embedded in a wide range of enterprise healthcare software.

Synthio Labs’ release does not break out an exact percentage for Azure, describing the result only as “fewer than half.” That is a meaningful gap in the public data. Anyone trying to judge exactly how Azure compares to ElevenLabs or Gemini on a like-for-like basis will need Synthio Labs, or Microsoft, to publish a precise figure.

How DOSE Scores a Pronunciation: The 0-to-5 Scale

DOSE grades each pronunciation on a scale of 0 to 5 against a set of verified reference pronunciations, with a score of 4 or higher counting as a pass. That threshold sets a fairly forgiving bar. A model does not need a perfect 5 to pass. It needs to land close enough to the reference pronunciation that a listener would recognize the word without hesitation. The fact that so many models still fell below that bar on newly approved names says less about an unreasonably strict rubric and more about how genuinely difficult novel pharmaceutical names are to synthesize correctly from text alone.

Drug names are deliberately engineered to be distinctive, often combining syllables that don’t map cleanly onto common English pronunciation patterns. That is by design, from a branding and drug-safety standpoint. It is also exactly the kind of input that trips up a text-to-speech model trained mostly on ordinary language, since the model has no dictionary entry, no common-word frequency signal, and no prior exposure to lean on.

DOSE Benchmark Results at a Glance

The table below summarizes the figures disclosed in Synthio Labs’ release, comparing established-name performance against newly-approved-name performance where both were published.

Model / MetricEstablished NamesNewly Approved NamesPoint Drop
ElevenLabs eleven_v393.0%67.1%-25.9 pts
Google Gemini TTS89.1%61.6%-27.5 pts
Microsoft Azure (generic/INN names)Not disclosedUnder 50%Not disclosed
All 9 models, overall pass range (all 274 names)63.1% – 80.3%
DOSE pass thresholdScore of 4 or higher on a 0-5 scale

DOSE Benchmark at a Glance: Scope and Structure

Beyond the headline scores, the mechanics of how DOSE was built help explain why the results carry weight. Synthio Labs designed the evaluation to mirror real clinical usage rather than isolated word tests, which is a meaningfully different approach than a simple pronunciation dictionary check.

MetricFigure
Benchmark nameDOSE (Drug-name Oral Synthesis Evaluation)
PublisherSynthio Labs
Announcement dateSeptember 17, 2026 (San Francisco)
Voice AI systems evaluated9
Total drug names tested274
Newly approved names in test set146
Test formatEach name read inside a real clinical sentence
Scoring scale0 to 5 against verified reference pronunciations
Pass thresholdScore of 4 or higher

Why Newly Approved Drug Names Break Voice AI Models

The pattern across every model in the DOSE release points to a training-data timing problem more than a raw capability problem. Large voice AI models learn pronunciation partly from patterns in written language and partly from any audio-text pairs they were trained on. Established drug names have had time to accumulate both. A name approved by regulators eight weeks before a model’s training cutoff has had neither, which forces the model to guess phonetically, treating a coined pharmaceutical name the way it would treat any unfamiliar string of letters.

That guessing problem does not go away with a bigger model or a longer context window. It is a data-freshness problem, and it will keep recurring every time a new batch of drugs clears approval, unless vendors build a pipeline to specifically ingest and verify pronunciations for newly approved pharmaceutical names before shipping them into consumer-facing products. Nothing in the DOSE release suggests any of the nine tested vendors currently does that as a matter of course.

The Clinical Stakes: Voice AI in Pharmacy and Telehealth

Drug-name confusion is not a new hazard invented by AI. Patient-safety groups have tracked look-alike and sound-alike drug name errors for decades, and organizations such as the Institute for Safe Medication Practices maintain running lists of medication names that clinicians confuse in speech and in writing. What DOSE adds to that long-standing problem is a machine in the loop that can introduce its own pronunciation errors at scale, across thousands of calls, texts, and reminders a day, without a human catching the mistake in the moment.

A pharmacist who mispronounces a drug name on a call is one error. A voice AI system that mispronounces the same drug name across every automated refill reminder it sends that week is a systemic one, and that risk only grows as voice assistants push into always-listening consumer hardware like AI glasses with built-in voice transcription. That distinction is exactly why a benchmark like DOSE matters more for voice AI than it would for, say, a chatbot that only outputs text. Text keeps the correct spelling on screen even if a reader mispronounces it privately. Audio removes that safety net entirely, particularly for patients who rely on voice output because they cannot easily read a screen.

Historical Context: Drug-Name Confusion Predates AI by Decades

Pharmaceutical naming has always fought an uphill battle against confusion. Regulators and industry groups have run formal programs for years to catch dangerous name collisions before a drug ever reaches a pharmacy shelf, which is part of why the WHO’s International Nonproprietary Names system exists in the first place: to give every drug a standardized generic name that reduces ambiguity across languages and markets. Voice AI did not create the underlying naming complexity. It inherited it, then added a new failure point on top of it, one that did not exist when the original name-confusion safeguards were designed decades before consumer text-to-speech was viable at scale.

What is genuinely new is the speed at which that inherited complexity now meets production software. A drug name that clinicians might spend months getting comfortable pronouncing can be pushed into a voice assistant’s vocabulary the same week it clears regulatory approval, with no equivalent adjustment period built in for the model.

How DOSE Compares to Other AI Safety and Capability Benchmarks

Most of the AI benchmarks that dominate headlines this year measure something broad: reasoning, coding, multi-turn dialogue, or raw factual recall. Model makers have leaned hard into that framing throughout 2026, with releases like GPT-6 Astra and open-weight contenders such as Qwen3.8-Max and GLM-5.2 all competing on general capability leaderboards. DOSE takes the opposite approach: it is deliberately narrow, testing exactly one skill, in exactly one domain, under conditions chosen to resemble real deployment rather than a lab-friendly word list.

That narrowness is a feature, not a limitation, for the audience DOSE is aimed at. A hospital IT buyer evaluating text-to-speech vendors does not need to know how a model performs on abstract reasoning puzzles. They need to know whether it will correctly say the name of the medication a patient is supposed to take. Domain-specific benchmarks like DOSE are likely to multiply across other high-stakes categories, following the same logic that led to specialized evaluations for legal, financial, and now pharmaceutical language.

Market Impact: What This Means for ElevenLabs, Google, and Microsoft

None of the three named companies has issued a public statement responding directly to the DOSE figures as of this writing. That silence is itself informative. Voice AI vendors have generally been happy to tout benchmark wins. A result this specific, and this unflattering on a safety-adjacent dimension, tends to get addressed quietly through an engineering fix rather than a press statement.

The commercial exposure differs by vendor. ElevenLabs sells voice synthesis as its core product, which means any credible, repeatable gap in pronunciation accuracy cuts directly at its main value proposition, especially with healthcare and pharma clients evaluating vendors. Google and Microsoft sell voice as one feature inside much larger platforms, Gemini and Azure respectively, which cushions the reputational hit somewhat but also means the fix has to travel through a bigger, slower release pipeline. Azure’s sub-50% result on generic drug names is arguably the most commercially sensitive of the three, given how deeply Azure AI Speech is already embedded in enterprise healthcare software that hospitals and clinics have already procured and deployed.

Competitive Comparison: Nine Models, One Weak Spot

Synthio Labs tested nine text-to-speech models in total, but the release names results for only three: ElevenLabs, Google, and Microsoft. The remaining six systems in the evaluation are not identified in the public release, which limits how precisely outside observers can rank the full field. What the release does make clear is that the overall pass range across all nine, 63.1% to 80.3%, sits well above any of the three named newly-approved-name scores, which means the drop on new drugs was not confined to the models singled out by name. It is very likely a shared weakness across most, if not all, of the systems tested.

That shared weakness is the real headline, arguably more than any single vendor’s individual score. A problem that only affected one company would be a company-specific engineering gap. A problem that shows up across a 63.1%-to-80.3% range of a nine-model field, with every disclosed model falling further on newly approved names, looks more like an industry-wide blind spot in how text-to-speech systems are trained and evaluated before shipping.

Predictions: Where Voice AI Pronunciation Testing Goes From Here

A handful of outcomes look likely in the months following the DOSE release, based on how similar narrow safety benchmarks have played out in other AI categories this year.

  • Expect at least one of the three named vendors, ElevenLabs, Google, or Microsoft, to publish a follow-up statement or updated model addressing newly approved drug names within the next quarter, given the specificity and public visibility of the DOSE numbers.
  • Expect Synthio Labs to expand DOSE into a recurring benchmark rather than a one-off release, tracking each new wave of drug approvals against the same nine-model (or larger) field, since a single snapshot has limited shelf life once new drugs enter the market.
  • Expect healthcare-sector buyers evaluating voice AI vendors for pharmacy, telehealth, or medication-reminder use cases to start asking for DOSE-style pronunciation scores as part of procurement, the same way accessibility and uptime scores already factor into enterprise voice AI contracts.
  • Expect other AI safety researchers to apply the same narrow-benchmark approach to adjacent domains where mispronunciation carries real risk, such as legal case citations, financial instrument names, or chemical compound names outside pharmaceuticals.
  • Expect scrutiny of the “one in three” headline figure to grow, given that Google’s own disclosed number in the release describes a worse failure rate than that framing suggests. A more precise accounting of all nine models, not just the three named, would settle exactly how representative that headline number really is.

What Voice AI Vendors Should Do Next

The most direct fix available to any of the nine vendors is not a bigger model. It is a faster, more disciplined data pipeline: a process that flags newly approved drug names as soon as they clear regulatory review and pushes verified reference pronunciations into the model’s training or fine-tuning data before that model ships into any pharmacy, telehealth, or medication-reminder product. That is closer to a content-operations problem than a research problem, which is arguably good news, since it does not require a breakthrough to fix, just sustained attention to a narrow but consequential slice of vocabulary.

In the meantime, healthcare organizations deploying any of these systems have a more immediate, low-tech option: keep a human review step in place for any voice AI output involving newly approved medications, at least until vendors can show a DOSE score, or something like it, that clears the same bar clinicians are already held to.

Frequently Asked Questions

What is the DOSE benchmark?
DOSE, short for Drug-name Oral Synthesis Evaluation, is a benchmark published by Synthio Labs on September 17, 2026, that tests how accurately text-to-speech and voice AI models pronounce drug names. Synthio Labs describes it as the first public benchmark focused specifically on this task.

How many models and drug names did DOSE test?
The benchmark evaluated nine voice AI models against 274 drug names, of which 146 were newly approved medications. Each name was tested inside a real clinical sentence rather than as an isolated word.

Which specific companies were named in the results?
The release names three vendors directly: ElevenLabs (eleven_v3), Google (Gemini TTS), and Microsoft (Azure). The remaining six models tested were not identified by name in the public release.

How much did accuracy drop on newly approved drug names?
ElevenLabs’ eleven_v3 dropped from 93.0% on established names to 67.1% on newly approved names. Google’s Gemini TTS dropped from 89.1% to 61.6%. Microsoft Azure passed fewer than half of generic (INN) drug names, according to the release.

Is the “one in three” figure accurate for every model?
It is a reasonably close description of ElevenLabs’ newly-approved failure rate, which works out to roughly 33%. It understates Google’s failure rate, which works out to roughly 38%, or nearly two in five. The claim as reported comes from Synthio Labs’ own framing of the release.

How does DOSE score a pronunciation as a pass or fail?
DOSE grades each pronunciation on a 0-to-5 scale against verified reference pronunciations. A score of 4 or higher counts as a pass.

Why do newly approved drugs cause more pronunciation errors than established ones?
Newly approved drug names have not had time to appear in training data, audio-text pairs, or common usage the way established names have, forcing voice AI models to guess phonetically from spelling alone, often with less accurate results.

Have ElevenLabs, Google, or Microsoft responded to the DOSE results?
As of this writing, none of the three named companies has issued a public statement responding to the DOSE benchmark results.