Three companies now sell the same basic product: on-demand access to frontier AI models through a cloud API. Amazon Bedrock, Microsoft’s Azure AI Foundry, and Google’s Vertex AI all promise the same pitch: pick a model, call an endpoint, pay per token. Under that shared pitch, the platforms diverge sharply on price, model selection, and how deep the tooling goes once an application moves past the demo stage.
By September 2026, the three platforms have also settled into distinct identities. Bedrock leans on breadth, offering more than 100 models from 18-plus providers through one API, according to AWS’s own Bedrock customer page. Azure AI Foundry leans on catalog scale, with Microsoft describing more than 10,000 models and roughly 50 new additions a month. Vertex AI leans on a tighter, more curated Model Garden, now past 200 models, with the deepest first-party integration for Google’s own Gemini family. None of the three is simply “the AI cloud” anymore. Each is a different bet on what enterprise AI infrastructure should look like.
This comparison walks through pricing for each platform’s current flagship model, catalog size, benchmark scores from independent trackers, uptime guarantees, real customer deployments, and the practical work of migrating a workload from one platform to another. The goal is a buying decision you can defend with numbers, not vendor marketing copy.
What Amazon Bedrock, Azure AI Foundry, and Vertex AI Actually Are
Amazon Bedrock is AWS’s managed layer for calling third-party and Amazon-built foundation models without provisioning GPU infrastructure yourself. It sits inside the AWS console next to Lambda, S3, and DynamoDB, so IAM policies, VPC endpoints, and CloudWatch logging all apply the same way they would to any other AWS service. Bedrock added Agents and AgentCore for building autonomous workflows, and it now includes web search grounding as a generally available feature, a capability AWS shipped alongside expanded Gemini access earlier in 2026.
Azure AI Foundry, the platform Microsoft previously split across Azure OpenAI Service and Azure Machine Learning, consolidated into one branded product and kept expanding through 2026. It is the default route for GPT-family models on Azure, but it also hosts Anthropic Claude, Meta Llama, Mistral, Cohere, xAI Grok, and NVIDIA-optimized models, all under one billing and governance layer. Microsoft’s pitch is portability of Provisioned Throughput Units (PTUs) across models, so a customer who reserves capacity is not locked to one specific model version.
Vertex AI is Google Cloud’s end-to-end machine learning platform, with Model Garden as the piece most directly comparable to Bedrock and Foundry. Where Bedrock and Foundry position themselves as neutral multi-vendor marketplaces, Vertex AI is still, first and foremost, the fastest and cheapest route to Google’s own Gemini models. Third-party models exist in Model Garden, including Anthropic Claude and Llama, but Google’s own model releases get first access to platform features like caching discounts and extended context windows.
Full Specs Comparison: Bedrock vs Azure AI Foundry vs Vertex AI
The table below lines up the core platform specs as of mid-September 2026. Model catalog counts come from each vendor’s own published figures; SLA figures come from each platform’s published service commitments.
| Spec | Amazon Bedrock | Azure AI Foundry | Google Vertex AI |
|---|---|---|---|
| Model catalog size | 100+ models, 18+ providers | 10,000+ models (Microsoft cites ~50 new/month) | 200+ models (Model Garden) |
| Current flagship reasoning model | Claude Opus 4.6 | GPT-6 Astra | Gemini 3.8 Flash (GA); Gemini 3.1 Pro (preview) |
| Flagship input pricing (per 1M tokens) | $5.00 | $10.00 (short context) / $20.00 (long context) | $0.75 |
| Flagship output pricing (per 1M tokens) | $25.00 | $50.00 (short) / $75.00 (long) | $3.75 |
| Max context window (flagship) | Not published by AWS for Opus 4.6 | Short/long context tiers, no published numeric cap | 1,048,576 input tokens / 65,536 output tokens |
| Uptime SLA | 99.9% monthly (invocation API) | 99.9% monthly (inherits Azure platform SLA) | 99.5% (Gemini Enterprise Agent Platform online inference); 99.9% for several other Vertex AI services |
| Native agent framework | Bedrock Agents / AgentCore | Foundry Agent Service, multi-agent orchestration | Gemini Enterprise Agent Platform |
| Fine-tuning support | Yes, per-model availability varies | Yes, with separate training + hosting billing | Yes, integrated into Model Garden deployment flow |
| Provisioned/reserved capacity | Provisioned throughput, priced per model unit/hour | Provisioned Throughput Units (PTUs), portable across models | Committed use discounts on Vertex AI compute |
| Batch pricing discount | ~50% off on-demand rate | Available via Foundry batch API | 50% off on-demand (batch input/output) |
| Cache read discount | $0.50 per 1M tokens (Opus 4.6) | $1.00-$2.00 per 1M tokens depending on context tier | Cached input tier available on Flash models |
| Primary console integration | AWS Console, IAM, VPC endpoints | Azure Portal, Entra ID, Azure Monitor | Google Cloud Console, IAM, VPC Service Controls |
Two things stand out immediately. First, the model catalog sizes are not comparable in any straightforward way: Azure’s “10,000+ models” figure includes many small, narrowly-scoped open-weight variants and fine-tuned derivatives, while Bedrock and Vertex AI report tighter, more curated numbers. Second, the pricing gap on flagship models is large. Gemini 3.8 Flash’s $0.75/$3.75 per-million-token rate is roughly one-seventh of Claude Opus 4.6’s on Bedrock and one-thirteenth of GPT-6 Astra’s short-context rate on Azure. That gap needs context, though, because Flash is Google’s fast, cheap tier, not a like-for-like reasoning competitor to Opus or Astra. Google’s closest reasoning-tier answer, Gemini 3.1 Pro, remains in preview with no published per-token price as of this writing.
Pricing Breakdown: Per-Token Costs Across the Three Platforms
Raw per-token pricing tells only part of the story, since batch discounts, cache pricing, and provisioned throughput all change the effective cost at scale. The table below covers standard on-demand, global-tier pricing for each platform’s current top-line model, plus a mid-tier reference point so the comparison isn’t skewed entirely by flagship-vs-flagship math.
| Platform / Model | Input ($/1M tokens) | Output ($/1M tokens) | Batch input | Batch output | Cache write |
|---|---|---|---|---|---|
| Bedrock: Claude Opus 4.6 (Standard Global) | $5.00 | $25.00 | $2.50 | $12.50 | $6.25 (5-min) / $10.00 (1-hr) |
| Bedrock: Claude 3.5 Sonnet (legacy reference) | $3.00 | $15.00 | $1.50 | $7.50 | $7.50 |
| Azure AI Foundry: GPT-6 Astra (short context, Global Standard) | $10.00 | $50.00 | Not itemized separately in Foundry pricing table | – | $12.50 |
| Azure AI Foundry: GPT-6 Astra (long context, Global Standard) | $20.00 | $75.00 | – | – | $25.00 |
| Azure AI Foundry: GPT-4o (legacy reference) | $2.50 | $10.00 | – | – | – |
| Vertex AI: Gemini 3.8 Flash | $0.75 | $3.75 | ~50% of on-demand | ~50% of on-demand | Cached input tier available |
| Vertex AI: Gemini 1.5 Pro (legacy reference) | $1.25 | $5.00 | $0.625 | $2.50 | – |
A practical example makes the gap concrete. A workload that processes 10 million input tokens and 2 million output tokens a day costs roughly $100 a day on Claude Opus 4.6 through Bedrock (10M input at $5/1M = $50, plus 2M output at $25/1M = $50). Run the same volume through GPT-6 Astra’s short-context tier and the bill rises to $200 a day (10M input at $10/1M = $100, plus 2M output at $50/1M = $100). Run it through Gemini 3.8 Flash and it drops to about $15 a day (10M input at $0.75/1M = $7.50, plus 2M output at $3.75/1M = $7.50). That is not a rounding difference. It is the reason FinOps teams increasingly route high-volume, latency-tolerant workloads to Flash-tier models and reserve flagship reasoning models for the smaller slice of requests that actually need them.
Benchmark Scores: How the Flagship Models Actually Perform
Pricing only matters relative to what a model can do. Three independent benchmark trackers, Vals AI, BenchLM, and Artificial Analysis, published scores for the current flagship models across these platforms through August and September 2026.
| Model | Platform | SWE-bench Verified | MMLU-Pro | Other benchmark | Source |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic API (Bedrock catalog lists Opus 4.6) | 97.00% | 91.59% | GPQA Diamond, MMMU top-ranked | Vals AI |
| Claude Fable 5.1 | Bedrock + Anthropic API | Not separately listed on SWE-bench Verified | 92.4% | GPQA Diamond 95.5% | BenchLM / Vals AI aggregation |
| GPT-6 Astra | Azure AI Foundry | Close behind Opus 5 on Vals SWE-bench leaderboard | Not itemized in this run | Terminal-Bench v4.0: 59%; AutomationBench-AA: 69%; Coding Agent Index: 67 | Artificial Analysis |
| GPT-5.6 Sol | Azure AI Foundry | 96.2% | Not itemized in this run | Artificial Analysis Intelligence Index: 58.9% (top of that index) | BenchLM |
| Gemini 3.8 Flash | Vertex AI | 80.0% | 90.2% | GPQA Diamond 94.4%; Vals Index 62.25% (7th of 57 models) | Vals AI / BenchLM |
The pattern that emerges is not “one platform wins.” Anthropic’s Claude Opus 5 sits at the top of the SWE-bench Verified leaderboard at 97.00%, per Vals AI’s benchmark tracker, but Bedrock’s own model cards, as documented in AWS’s Anthropic model listings, still reference Claude Opus 4.6 as the flagship Opus-tier offering, which means Bedrock customers may be running a generation behind Anthropic’s own top-scoring release depending on which model ID they deploy. GPT-6 Astra, meanwhile, leads on agent-oriented benchmarks like AutomationBench-AA and Terminal-Bench v4.0, reflecting OpenAI’s focus on long-horizon coding and tool-use tasks. Gemini 3.8 Flash trails both on raw SWE-bench score at 80.0%, but it was never built to compete there. It is priced and designed as the cheap, fast, high-volume tier, and on that axis it delivers a defensible score for a fraction of the cost.
Model Catalog Depth: Breadth vs Curation
Catalog size sounds like a simple number to compare, but the three vendors count differently. Amazon Bedrock’s own materials describe the platform growing from six models at 2023 launch to more than 100 today, spanning 18-plus providers including Anthropic, Meta, Mistral, Cohere, Stability, and Amazon’s own Titan and Nova families. That is a curated list: each model gets its own model card, its own pricing row, and its own inference profile.
Azure AI Foundry’s catalog is an order of magnitude larger by Microsoft’s own count, described in Foundry’s model overview documentation as exceeding 10,000 models with roughly 50 new additions monthly. That figure includes fine-tuned variants, smaller open-weight checkpoints, and models contributed by third parties like NVIDIA and Fireworks, not just headline frontier releases. For a team searching for a specific narrow-domain open-weight model, that depth matters. For a team deciding between three or four frontier reasoning models, it is mostly noise.
Vertex AI’s Model Garden sits in between at 200-plus models, per Google Cloud’s own Next 2026 announcements, spanning Google’s Gemini and Gemma families, Anthropic’s Claude, and open-weight models like Llama. Google’s positioning is less “biggest catalog” and more “best-integrated catalog for our own models,” which shows up in how quickly new Gemini releases get platform-level features like caching and audio input support. Google added new SKUs for Gemini 3 Flash Audio Input to its Vertex AI SKU groups within days of that capability shipping, a turnaround that neither Bedrock nor Foundry can match for Google’s own models by definition.
Developer Experience: API Design, SDKs, and Onboarding
Pricing and benchmarks get most of the attention in a platform comparison, but the day-to-day experience of writing code against these APIs shapes adoption just as much. Amazon Bedrock’s InvokeModel and Converse APIs follow AWS’s long-standing SDK conventions: boto3 for Python, the AWS SDK for Java, and consistent request signing through IAM. That consistency is a strength for teams already writing AWS-native code, since a Bedrock call looks and behaves like any other AWS service call, down to retry logic and CloudWatch metrics. It is a weakness for teams new to AWS, since the same IAM permission model that makes Bedrock secure also makes a first integration slower than a plain API-key setup.
Azure AI Foundry’s SDK, built around the Foundry Models client library, mirrors the OpenAI Python and JavaScript SDKs closely enough that teams migrating from a direct OpenAI API integration typically report a lighter lift than moving to either Bedrock or Vertex AI. Microsoft has leaned into this deliberately, keeping request and response schemas close to OpenAI’s own format even for non-OpenAI models hosted in the same catalog. The tradeoff is that Foundry’s unified governance layer, while powerful for large organizations managing dozens of deployed models, adds setup overhead (resource groups, deployment names, region selection) that a solo developer testing a prototype doesn’t need.
Vertex AI’s generateContent and streamGenerateContent methods, part of the Gemini API surface, tend to draw the most praise from developers for raw simplicity, particularly for multimodal inputs. Sending an image, a video clip, and a text prompt in a single request is more straightforward on Vertex AI than on the other two platforms, a reflection of Gemini’s natively multimodal training versus the more bolted-on multimodal support in some competing models. The cost of that simplicity shows up in tooling maturity elsewhere: Vertex AI’s agent framework and enterprise governance features are newer than Bedrock’s or Foundry’s, so documentation and community examples are thinner for advanced agent orchestration patterns compared with basic inference calls.
Enterprise Case Studies: Who Is Actually Running What
Vendor pricing pages are marketing. Named customer deployments are closer to evidence. Here are five documented, publicly disclosed 2026 examples across the three platforms.
- AstraZeneca on Amazon Bedrock — AWS’s Bedrock customers page lists AstraZeneca using Bedrock Agents to accelerate drug development decisions and research insights.
- DoorDash on Amazon Bedrock — An AWS solutions case study describes DoorDash building a generative AI self-service contact center using Amazon Bedrock, Amazon Connect, and Anthropic’s Claude models.
- Domcura on Azure AI Foundry — The German insurer built an AI claims processing platform called “Kim” on Microsoft Foundry and Azure, cutting payout processing time to about 10 minutes and reducing operational costs by 50%, according to Technology Record’s reporting.
- Mercedes-Benz and Wells Fargo on Vertex AI — Coverage of Google Cloud Next 2026 cited more than 47,000 enterprise customers deploying agents on Vertex AI, naming Mercedes-Benz’s engineering agents and Wells Fargo’s compliance agents as marquee examples.
- Proton on Vertex AI — A Nucleus Research ROI case study found Proton achieved a 232% return on investment and a 9.5-month payback period after adopting Vertex AI as its machine learning and generative AI platform.
The pattern across these five: Bedrock deployments skew toward regulated, data-sensitive workloads already living inside AWS (pharma, logistics, contact centers). Azure Foundry deployments skew toward European enterprise customers already standardized on Microsoft’s identity and compliance stack. Vertex AI deployments skew toward organizations that want the newest Gemini capabilities fastest, especially in engineering and large-scale agent orchestration. None of this is universal, but it tracks with how each cloud markets itself.
Market Position and Cloud Revenue Context
None of these AI platforms exists in isolation from the parent cloud’s broader financial position. Microsoft’s own investor materials for fiscal year 2026 disclosed Microsoft Cloud revenue of $54.5 billion in fiscal Q3, up 29% year over year, with Azure and other cloud services growing roughly 40%. The same FY26 Q3 earnings materials stated that Microsoft’s AI business annual revenue run rate surpassed $37 billion that quarter, growing 123% year over year, a figure that spans Azure AI Foundry, Copilot, and GitHub Copilot combined rather than Foundry alone.
AWS and Google Cloud report their AI-specific revenue less granularly in public filings, folding it into broader cloud infrastructure segments. Independent market trackers put overall cloud infrastructure share at roughly AWS 31%, Azure 24%, and Google Cloud 12% for Q2 2026, based on a report from 16IDC, an industry analytics firm distinct from the more widely cited IDC research group. Those figures describe total cloud infrastructure spend, not AI-platform-specific share, so they should be read as context for the parent companies’ scale rather than a direct read on Bedrock-versus-Foundry-versus-Vertex adoption.
That scale gap matters for buyers in one practical way: capacity availability. When a hyperscaler is growing its AI infrastructure revenue 40% or more year over year, GPU and provisioned-throughput capacity for the newest flagship models can sell out or queue in high-demand regions, a pattern several 2026 pricing trackers have noted for Blackwell-class GPU capacity blocks across all three clouds. Enterprises planning a large provisioned-throughput commitment on any of the three platforms should treat capacity confirmation, not just pricing, as part of the procurement process, particularly for launch-week access to a model like GPT-6 Astra or Gemini 3.1 Pro.
Uptime, SLA Terms, and Reliability Track Record
All three vendors publish formal SLA commitments, and all three use a similar tiered-credit structure: partial service credit for uptime below the target, larger credits the further uptime falls. Amazon Bedrock’s SLA commits to 99.9% monthly uptime for the Bedrock invocation API, with credit tiers of 10% between 99.0% and 99.9% uptime, 25% between 95.0% and 99.0%, and 100% below 95.0%. Azure OpenAI inherits the broader Azure platform SLA, also set at 99.9% monthly uptime with matching credit tiers, though the SLA explicitly covers endpoint availability only and excludes latency or throughput degradation.
Google’s SLA structure for Vertex AI is more segmented by service. The Gemini Enterprise Agent Platform’s online inference methods (generateContent and streamGenerateContent) carry a 99.5% monthly uptime target, slightly below the other two platforms’ 99.9%, with credit tiers of 10% between 99.0% and 99.5%, 25% between 95.0% and 99.0%, and 50% below 95.0%. Other Vertex AI services, including AutoML Tabular and Image online prediction on two or more nodes, carry the higher 99.9% target. In practice, this means the specific Vertex AI feature you use determines which SLA tier applies, a wrinkle that Bedrock and Azure OpenAI customers do not have to think about as much since their flagship inference APIs sit under one uniform number.
Security, Compliance, and Governance Features
Each platform inherits its parent cloud’s compliance posture rather than building a separate certification stack for AI specifically. Bedrock runs under AWS’s existing IAM, VPC endpoint, and CloudTrail logging infrastructure, plus a dedicated Guardrails feature for content filtering, billed separately from base inference. Azure AI Foundry runs under Azure’s Entra ID and Azure Policy framework, with the added wrinkle of regional data-zone pricing: Microsoft introduced regional multipliers effective September 1, 2026, charging 10% to 50% above global rates depending on the deployment region, a change that affects EU, APAC, and other non-US data zones differently.
Vertex AI runs under Google Cloud’s IAM and VPC Service Controls, with the Gemini Enterprise Agent Platform adding its own governance layer for agent permissions and tool access. For teams already committed to one cloud’s identity provider, sticking with that cloud’s AI platform avoids a second layer of access control to manage. For teams genuinely undecided, the compliance differences between the three are smaller than the pricing and model-access differences, and shouldn’t be the deciding factor on their own.
Data residency is the one compliance-adjacent area where the platforms genuinely diverge in cost, not just policy. Azure’s September 2026 regional pricing update is the clearest example: an EU customer running GPT-6 Astra in the EU Data Zone now pays roughly 20% above global list price, on top of the 9% increase already applied to that zone earlier in the year. Bedrock and Vertex AI both offer regional deployment options for data residency, but neither has published a comparable public multiplier structure tying region choice directly to per-token cost. For a European enterprise weighing all three platforms, that difference alone can outweigh the base per-token pricing gap discussed earlier in this comparison, and it’s worth confirming with each vendor’s account team before signing a large commitment.
Best Use Cases for Each Platform
Matching workload to platform matters more than picking an abstract winner. Based on the pricing, benchmark, and case-study data above, here is where each platform tends to fit best. Note that these are starting points, not fixed rules: a team running a mixed workload of customer support automation and internal coding agents may reasonably end up using two of these three platforms simultaneously, and increasingly does, given how far apart the flagship pricing sits.
- High-volume customer support automation on an existing AWS stack — Amazon Bedrock, especially paired with Claude via Bedrock Agents, as demonstrated by DoorDash’s contact center build.
- Regulated industries already standardized on Microsoft identity and compliance tooling — Azure AI Foundry, particularly for European enterprises where Microsoft’s data-zone options and existing Azure AD integration reduce migration friction.
- Large-scale agent orchestration across engineering, compliance, and HR functions — Vertex AI’s Gemini Enterprise Agent Platform, given the 47,000-plus enterprise agent deployments Google reported at Cloud Next 2026.
- Cost-sensitive, high-throughput inference where model ceiling matters less than price per call — Vertex AI’s Gemini 3.8 Flash tier, at roughly one-seventh the per-token cost of the flagship models on Bedrock or Foundry.
- Maximum coding-agent and long-horizon tool-use performance — Azure AI Foundry’s GPT-6 Astra, which leads on Terminal-Bench v4.0 and AutomationBench-AA among the models covered here.
- Frontier reasoning tasks requiring the top published SWE-bench Verified score — Anthropic’s Claude Opus 5, though buyers should confirm which exact Claude model ID is deployed on Bedrock, since AWS’s own documentation currently references Opus 4.6 rather than Opus 5 as the Bedrock flagship.
Migration Guide: Moving a Workload Between Platforms
Switching AI platforms is rarely a full rewrite, but it is never a one-line config change either. Here is the realistic path for moving an existing generative AI workload from one of these three platforms to another.
- Audit current API calls and abstract the model-invocation layer. If your application calls Bedrock’s InvokeModel API, Azure’s Foundry SDK, or Vertex AI’s generateContent method directly throughout the codebase, introduce a thin wrapper first. This is the single highest-leverage step and should happen regardless of whether a migration is imminent.
- Map prompt formats between providers. Claude, GPT, and Gemini each expect slightly different system-prompt structures, message role conventions, and function-calling schemas. Budget real testing time here; this is where most migration bugs surface, not in the billing setup.
- Re-benchmark on your own evaluation set, not published leaderboards. SWE-bench and MMLU-Pro scores tell you about general capability, not your specific prompts. Run your existing eval suite against the new target model before committing.
- Recalculate cost at your actual token volume. Use your last 30 days of production input/output token counts against each platform’s pricing table above, including batch and cache discounts if your workload qualifies.
- Migrate fine-tuned models or retrain equivalents. Fine-tuned weights generally do not transfer across providers. If your workload depends on a fine-tuned model, budget for a retraining pass on the new platform’s fine-tuning pipeline.
- Update IAM and identity/network configuration. Moving from Bedrock to Foundry means shifting from IAM roles and VPC endpoints to Entra ID and Azure Private Link (or the reverse). This is usually a bigger lift than the API changes themselves in regulated environments.
- Rebuild agent tool definitions. Bedrock Agents, Foundry Agent Service, and the Gemini Enterprise Agent Platform each define tool and function schemas differently. Agent-heavy workloads should treat this as its own migration phase, separate from basic inference calls.
- Run a shadow-traffic period before cutover. Route a percentage of real traffic to the new platform while keeping the old one live, comparing latency, error rates, and output quality before fully switching.
- Decommission the old platform’s provisioned capacity last. If you were running provisioned throughput on Bedrock or PTUs on Azure, cancel or downsize that reserved capacity only after the shadow-traffic period confirms the new platform is stable, to avoid paying for both simultaneously any longer than necessary.
Pros and Cons: Amazon Bedrock
Bedrock’s strongest case is for teams already deep in the AWS ecosystem who want multi-vendor model access without leaving their existing IAM and networking setup.
| Pros | Cons |
|---|---|
| Deep native integration with existing AWS services (Lambda, DynamoDB, S3) | Model catalog can lag behind a provider’s own latest release (Opus 4.6 listed vs Opus 5 on leaderboards) |
| Broad multi-vendor model selection (18+ providers) through one API | Pricing sits mid-pack, cheaper than Azure’s flagship but pricier than Vertex AI’s Flash tier |
| Mature Guardrails and Agents/AgentCore tooling | Context window figures for flagship models are not always clearly published |
| Straightforward 99.9% SLA on the core invocation API | Batch and cache pricing details vary noticeably by exact model ID |
Pros and Cons: Azure AI Foundry
Foundry’s strongest case is for enterprises already standardized on Microsoft identity and governance tooling, particularly in regulated European markets.
| Pros | Cons |
|---|---|
| Largest published model catalog (10,000+) with fastest new-model cadence | Highest flagship pricing of the three ($10-$20 input / $50-$75 output per 1M tokens for GPT-6 Astra) |
| GPT-6 Astra leads on agentic/coding-agent benchmarks like AutomationBench-AA | New regional pricing multipliers (up to 50% above global rate) complicate cost forecasting outside the US |
| PTU capacity is portable across models, reducing reservation lock-in | SLA explicitly excludes latency/throughput guarantees, only covers endpoint availability |
| Deep Entra ID and Azure Policy integration for enterprise governance | Catalog size includes many narrow variants, making genuine model comparison harder |
Pros and Cons: Google Vertex AI
Vertex AI’s strongest case is cost-sensitive, high-throughput workloads and organizations that want the fastest access to new Gemini capabilities.
| Pros | Cons |
|---|---|
| Lowest flagship-tier pricing by a wide margin ($0.75/$3.75 per 1M tokens on Gemini 3.8 Flash) | Lowest published SLA of the three for online inference (99.5% vs 99.9%) |
| Largest 1M-token context window among the models compared here | Top reasoning-tier model, Gemini 3.1 Pro, remains in preview with no published pricing |
| Fastest rollout of new Gemini features to the platform (audio input SKUs shipped within days) | Third-party model selection in Model Garden is narrower than Bedrock or Foundry |
| Strong documented enterprise agent adoption (47,000+ customers per Google’s Cloud Next 2026 figures) | Best pricing is tied to the Flash tier, not the top reasoning tier, so cheapest and most capable aren’t the same SKU |
The Verdict: Which Platform Should You Actually Pick
There is no single winner here, and any comparison that claims otherwise is selling something. The data points to three different correct answers depending on what you are optimizing for.
If your priority is raw model capability on coding and reasoning tasks and you can tolerate premium pricing, Azure AI Foundry’s GPT-6 Astra currently leads on agent-oriented benchmarks like Terminal-Bench v4.0 (59%) and AutomationBench-AA (69%), per Artificial Analysis’s own benchmarking. If your priority is minimizing per-token cost at scale, Vertex AI’s Gemini 3.8 Flash is not close: at $0.75 input and $3.75 output per million tokens, it costs a fraction of the other two flagships, while still scoring a respectable 90.2% on MMLU-Pro. If your priority is model breadth, vendor neutrality, and staying inside an existing AWS deployment, Amazon Bedrock’s 100-plus models across 18 providers remains the most flexible single API, even if its listed flagship, Opus 4.6, trails Anthropic’s own top-scoring release on independent leaderboards.
The more useful takeaway for most teams: this is not a single-platform decision anymore. The pricing gaps are wide enough, and the benchmark differences specific enough, that running a small workload split, cheap high-volume calls on Vertex AI’s Flash tier, premium reasoning calls on Bedrock or Foundry, is now a defensible architecture rather than unnecessary complexity. Multi-cloud AI routing has stopped being a hedge against vendor lock-in and started being a straightforward cost-optimization strategy.
Frequently Asked Questions
Is Amazon Bedrock cheaper than Azure AI Foundry?
For comparable flagship models, yes. Claude Opus 4.6 on Bedrock lists at $5.00 input and $25.00 output per million tokens, compared with GPT-6 Astra’s $10.00 input and $50.00 output for short context on Azure AI Foundry. Both are more expensive per token than Vertex AI’s Gemini 3.8 Flash, though Flash is a lighter-weight model tier, not a direct flagship-to-flagship comparison.
Which platform has the most AI models available?
By raw count, Azure AI Foundry, with Microsoft citing more than 10,000 models in its catalog. Amazon Bedrock offers more than 100 curated models from 18-plus providers, and Google Vertex AI’s Model Garden lists more than 200 models. The right comparison depends on whether you need catalog depth or a smaller, more curated selection.
Can I use Claude on Azure or Gemini on AWS?
Yes. Azure AI Foundry’s catalog includes Anthropic Claude models alongside GPT, and Google’s Model Garden includes third-party models including Claude. AWS Bedrock added expanded access to Google’s Gemini models as part of a 2026 platform update. None of the three clouds restricts itself entirely to its own first-party models anymore.
What is the uptime SLA for each platform?
Amazon Bedrock and Azure OpenAI/Foundry both commit to 99.9% monthly uptime for their core inference APIs. Google’s Gemini Enterprise Agent Platform online inference methods carry a 99.5% target, though other Vertex AI services carry a 99.9% target. All three use tiered service credits for uptime shortfalls rather than flat penalties.
Which platform is best for building AI agents?
All three now offer dedicated agent tooling: Bedrock Agents/AgentCore, Azure AI Foundry’s Agent Service, and the Gemini Enterprise Agent Platform. Google reported the largest publicly disclosed enterprise agent deployment figure, more than 47,000 customers, at Cloud Next 2026, though that figure reflects Google’s own reporting rather than an independent audit.
Do these platforms offer free tiers or trial credits?
Published documentation for all three platforms in 2026 focuses on paid, pay-as-you-go pricing rather than dedicated free-tier token allowances for enterprise generative AI usage. Teams evaluating any of the three should check current promotional credits directly with each vendor, since these change frequently and were not consistently documented in the pricing pages reviewed for this comparison.
Why does Bedrock list Claude Opus 4.6 instead of Claude Opus 5?
Cloud platforms do not always carry a model provider’s newest release on day one. Anthropic’s Claude Opus 5 tops independent SWE-bench Verified leaderboards at 97.00%, according to Vals AI, but AWS’s own Bedrock model card documentation, as of this comparison, still lists Claude Opus 4.6 as its flagship Opus-tier offering. Buyers who need the absolute newest Anthropic model should verify current Bedrock model availability directly before assuming leaderboard-topping models are already live on every platform.
How hard is it to switch between these three platforms?
Moving basic inference calls is manageable with an abstraction layer, but fine-tuned models generally don’t transfer, agent tool definitions need rebuilding, and identity and networking configuration (IAM vs Entra ID vs Google Cloud IAM) is usually the biggest hidden cost. Budget for a shadow-traffic testing period before fully cutting over any production workload.



