AI MODEL BRIEFUpdated 17 September 2026

Amazon Bedrock & Microsoft Foundry

Choosing models for
software development.

Model capabilities, cloud availability and token prices for planning and implementation in a coding agent.

01

Models and coding performance

Top-tier and everyday models for planning, implementation and review.

Model developerTop tierEveryday workhorseCoding performance
OpenAIUnited States
GPT-5.6 SolTop-tier coding reference
GPT-5.6 TerraGeneral coding at lower cost
SolSWE-Pro: 64.6% · TB 2.1: 88.8%
TerraSWE-Pro: 63.4% · TB 2.1: 87.4%
AnthropicUnited States
Claude Opus 5Complex implementation and review
Claude Sonnet 5Everyday agentic development
Opus 5SWE-Pro: 79.2% · TB 2.1: —
Sonnet 5SWE-Pro: 63.2% · TB 2.1: 80.4%
DeepSeekChina
V4-Pro-0813Current Pro snapshot · August 2026
V4.1 FlashLatest Flash · September 2026
Pro-0813SWE-Pro: — · TB 2.1: 87.9%
V4.1 FlashSWE-Pro: — · TB 2.1: 90.6%
Moonshot AIChina · Kimi
Kimi K3Flagship · 1M context
Kimi K2.7 CodeCoding specialist · 256K context
K3SWE-Pro: — · TB 2.1: 88.3%
K2.7 CodeSWE-Pro: — · TB 2.1: —Its model card reports other coding benchmarks.
AlibabaChina · Qwen
Qwen3.8-MaxLatest snapshot: 0902
Qwen3.8-FlashLower-cost 3.8 variant
Max · original releaseSWE-Pro: 67.7% · TB 2.1: 86.6%Not verified for 0902; SWE-Pro uses corrected tasks.
FlashSWE-Pro: 62.5% · TB 2.1: —
MiniMaxChina
MiniMax M3Latest flagship
MiniMax M3Also the economical choice
M3 · both rolesSWE-Pro: 59.0% · TB 2.1: 66.0%
Z.aiChina · GLM
GLM-5.3Flagship · August 2026
GLM-5.3-FlashLower-cost current variant
GLM-5.3SWE-Pro: — · TB 2.1: —Reports 28.3% on Terminal-Bench 3.0, a different test.
5.3-FlashSWE-Pro: — · TB 2.1: 84.3%

SWE-Pro = SWE-bench Pro, repository-level software tasks. TB 2.1 = Terminal-Bench 2.1, tasks performed through a terminal. Higher percentages indicate more tasks solved within each evaluation. A dash means no result was verified for that exact model and benchmark; it is not a zero score.

Results are reported by the model developers under different agent setups, reasoning effort and evaluation conditions, so they are not a controlled ranking. Sonnet uses max effort for SWE-Pro and xhigh for TB 2.1; Qwen Max uses corrected SWE-Pro tasks. Terminal-Bench 3.0 is not comparable to 2.1. Each model label in the performance column links to its source.

Premium frontier · OpenAI

GPT-6 Astra

For particularly demanding planning and development.

$10 input / $50 output

Per million origin API tokens. 2.5× Sol’s standard token rates. Model details

Premium frontier · Anthropic

Claude Fable 5.1

For complex, long-running coding work.

$10 input / $50 output

Per million origin API tokens. 2× Opus 5’s standard token rates. Model details

Astra and Fable are excluded from the savings baselines below. Their higher token rates can still be offset by fewer tokens, fewer retries or caching. Fable 5.1 cache reads cost $0.25 per million tokens.

02

Model versions by provider

Current origin models and the pay-per-token versions available through each cloud.

DeveloperLatest at the originAmazon BedrockMicrosoft Foundry
OpenAI
GPT-5.6 Sol / TerraCurrent practical pair
GPT-5.6 Sol / TerraSame model names; separate regional routing rulesCurrent family
GPT-5.6 Sol / TerraBoth: version 2026-07-09 · Azure DirectCurrent family
Anthropic
Opus 5 / Sonnet 5Current practical pair
Opus 5 / Sonnet 5Same model names; EU-local options existCurrent family
Opus 5 / Sonnet 5GA. v1: Anthropic infrastructure; v2: Azure infrastructure. Anthropic operates both.Current family
DeepSeek
V4-Pro-0813Current ProV4.1 FlashLatest Flash
V3.2Earlier than both current origin choicesEarlier release
V4 Pro2026-04-23; older than Pro-0813V4-Flash-07312026-07-31 · preview; earlier than V4.1Earlier snapshots
Moonshot / Kimi
K3 / K2.7 CodeFlagship / coding workhorse
K2.5Earlier generation; retired from the origin APIEarlier release
K3 / K2.7 CodeFireworks partner · pay per tokenK2.7 CodeAlso Azure Direct · 2026-06-12 · previewCurrent names
Alibaba / Qwen
3.8-Max / 3.8-FlashMax snapshot: qwen3.8-max-0902
Qwen3 Coder 480BLarge coding model; distinct from Qwen3.8Qwen3 Coder Next / 30BSmaller coding alternativesDifferent coding branch
No token offer verifiedQwen catalogue weights and provisioned offers do not establish a pay-per-token route.No PAYG verified
MiniMax
M3Latest model; also the workhorse
M2.5Earlier than M3Earlier release
M3Fireworks partner · 512K cloud context; origin supports 1MCurrent name
Z.ai / GLM
5.3 / 5.3-FlashFlagship / workhorse
GLM 5 / GLM 4.7 FlashBoth earlier than the current pairEarlier releases
GLM-5.3Fireworks partner · pay per token5.3-Flash not listedGLM-5.2-Fast is a different alternative.Flagship only

“Current” means the published model family or name matches, not that weights, context limits or serving behaviour are identical. Availability here is documented pay-per-token availability; account access and capacity were not tested.

Premium frontier availability

ModelAmazon BedrockMicrosoft Foundry
GPT-6 AstraGlobal routing
GPT-6 AstraVersion 2026-09-03 · Global Standard
Fable 5.1Global routing from European endpoints
Fable 5.1Version 1 · Anthropic-hosted preview

Newer Chinese models can also be purchased through AWS Marketplace. Fireworks AI offers current DeepSeek V4.1 Flash, V4 Pro-0813, Qwen3.8-Max, Kimi K3, MiniMax M3 and GLM-5.3/Flash through its external per-token API. These are separate from the native Bedrock catalogue. See the model-by-model prices.

03

European access and inference

The location of an endpoint can differ from the location where a model processes the request.

European endpoint

The request enters the service in Europe. A Global deployment can still process it elsewhere.

Processing stays in a defined area

In-region processing or a geographic deployment policy limits where inference runs. The exact area matters.

Amazon Bedrock

Available modelsEuropean endpointInference confined to the EU?
Opus 5 / Sonnet 5
Stockholm or Ireland
Yes, with in-region deploymentAvailable through in-region deployment and EU geographic profiles. Model routing
Astra / Sol / Terra
Yes, Global routing
No on the documented European routesAn EU endpoint does not limit their Global inference. OpenAI routing
Fable 5.1
Yes, Global routing
No on the documented European routesFable routing
DeepSeek V3.2 / Kimi K2.5
Stockholm
Yes, in-regionThese are older releases. DeepSeek · Kimi
MiniMax M2.5 / GLM 5 / GLM 4.7 Flash
Stockholm; more EU regions for some models
Yes, in-regionMiniMax · GLM 5 · GLM Flash
Qwen3 Coder 480B / 30B
Stockholm
Yes, in-regionCoder 480B · Coder 30B
Qwen3 Coder Next
London is documented; EU entries conflict
Not verified for the EUPricing lists EU locations that the model card does not confirm. London is outside the EU. Model card

Microsoft Foundry

Available modelsEuropean endpointInference confined to the EU?
GPT-5.6 Sol / Terra
EU Data Zone is listed
EU/EFTA area, not EU27-onlyCurrent deployment policy includes EFTA. Terra’s Standard availability and tariff are not consistently documented. Deployment policy
GPT-6 Astra
Yes, Global Standard
No EU Data Zone Standard listedRegional availability
Fable 5.1 / Opus 5 / Sonnet 5
Sweden Central · Global
No EU-confined offer verifiedThe European entry uses Global routing. Partner deployment table
DeepSeek V4-FlashOriginal 2026-04-23 version
EU Data Zone Standard
EU/EFTA area, not EU27-onlyThis is an older Flash snapshot. Regional availability
V4 Pro / Flash-0731 / K2.7 Code
Yes, Global Standard
No EU-confined route verifiedRefers to the Azure Direct models. Regional availability
Fireworks K3 / K2.7 Code / M3 / GLM-5.3
No European pay-per-token deployment listed
US Data Zone onlyThese Fireworks offers are excluded from the EU Data Boundary. Fireworks deployment table
Qwen3.8
No qualifying token offer verified
Not establishedA model-weight catalogue entry is not a serverless inference commitment.

Native Bedrock offers several older Chinese models with EU-local processing. Foundry offers more recent Kimi, MiniMax and GLM models through Fireworks, with US processing.

Fireworks purchased separately through AWS Marketplace has no published EU-only guarantee for default serverless inference. Its published restrictions are unrestricted or US; other locations require a separate offer. Fireworks residency.

This section describes inference processing. It does not establish where all logs, abuse-monitoring records or other service data are retained. AWS routing · Azure deployment types

04

Model prices by tier

Origin API prices, with Opus 5 and Sol as top-tier references and Sonnet 5 and Terra as everyday references.

Indicative token costs. These are candidates for the same development roles. No matched evaluation on Anritsu’s tasks establishes that they can replace Opus 5 or Sol at equal quality.

All rates below are USD per million tokens at the model developer’s API. The illustrative bill uses 1 million input + 200,000 output tokens. It holds token volume fixed; it is not a measured cost per completed task.

Top tier: Opus 5 and Sol as references

ModelInput / outputIllustrative token billSaving at equal token volume
$5 / $25
$10.00
Reference model
$4 / $20
$8.00
Reference model
$1.32 / $3.96
$2.11
79% less than Opus 574% less than Sol
$3 / $15
$6.00
40% less than Opus 525% less than Sol
$2 / $6
$3.20
68% less than Opus 560% less than Sol
$0.3 / $1.2
$0.54
95% less than Opus 593% less than Sol
$1.4 / $4.4
$2.28
77% less than Opus 572% less than Sol

At the same token budget, these Chinese candidates cost 25–93% less than Sol and 40–95% less than Opus 5. The unweighted average saving across the five Chinese developers is 72% versus Opus 5. Kimi K3 sits at the expensive end; MiniMax M3 at the inexpensive end.

Everyday work: Sonnet 5 and Terra as references

ModelInput / outputIllustrative token billSaving at equal token volume
$2 / $10
$4.00
Reference model
$2 / $12
$4.40
Reference model
$0.3 / $1.2
$0.54
86% less than Sonnet 588% less than Terra
$0.95 / $4
$1.75
56% less than Sonnet 560% less than Terra
$0.15 / $0.47
$0.24
94% less than Sonnet 594% less than Terra
$0.3 / $1.2
$0.54
86% less than Sonnet 588% less than Terra
$0.15 / $0.5
$0.25
94% less than Sonnet 594% less than Terra

For the everyday choices, the unweighted average saving across the five Chinese developers is 83% versus Sonnet 5, with individual savings of 56–94%. These are price comparisons within broad usage tiers, not measured equivalence in coding capability.

Basis: standard uncached rates, DeepSeek peak hours, Qwen Singapore International, and inputs within the lowest context tier. No batch discounts, taxes, tool charges or partner fees. Sol’s $4/$20 origin offer is stated through at least 21 November 2026. Sol terms

Billing assumptions

DeepSeek rates are half-price outside its published peak windows. Qwen Frankfurt Global is lower than Singapore International: Max $1.65/$4.951 and Flash $0.113/$0.382. A Frankfurt Global endpoint does not establish EU-only inference.

Large prompts can trigger higher tiers: OpenAI above 272K input tokens; MiniMax M3 above 512K. The table uses one million input tokens accumulated across requests within the base tier, not a single one-million-token prompt. Reasoning output, cache reads and writes, retries, and successful task completion all affect the actual bill. Different models can tokenize the same text differently.

Provider benchmarks use different harnesses, effort levels and task versions. They support a shortlist, but not a single cross-provider claim of equal coding capability. DeepSeek terms · Qwen pricing · MiniMax pricing

05

Origin and cloud prices

Rates depend on the model version, processing location and commercial offer.

Same published model names are compared below. USD per million input / output tokens. Standard Global rates unless a cell states US Data Zone; these are not EU-local quotes.

ModelOrigin APIAmazon BedrockMicrosoft Foundry
$4 / $20
$4 / $20Global
$2 / $12
$2 / $12Global
$2 / $12Global
$5 / $25
$5 / $25Global
$2 / $10
$2 / $10Global
$3 / $15
Not listed natively
$3.3 / $16.5Fireworks · US Data Zone · 10% higher
$0.3 / $1.2
Not listed natively
$0.33 / $1.32Fireworks · US Data Zone · 10% higher
$1.4 / $4.4
Not listed natively
Available via FireworksExact GLM-5.3 tariff not verified. Pricing

Bedrock EU-local Opus 5 costs $5.50/$27.50; Sonnet 5 costs $2.20/$11, 10% above the Global rates. Astra and Fable 5.1 cost $10/$50 on all three Global routes. Foundry CCU is token-derived usage billing, not payment for hosted compute. AWS tariffs · Foundry Claude billing

EU prices for earlier Bedrock models

Exact modelOrigin API todayBedrock in the EU
No longer served directlyOrigin release history
$0.74 / $2.22Stockholm · in-region
Discontinued on 31 AugustOrigin model catalogue
$0.72 / $3.6Stockholm · in-region
$0.45 / $1.8Stockholm · in-region
$0.3 / $1.2
$0.36 / $1.44Stockholm · in-region
$1 / $3.2
$1.2 / $3.84Stockholm · in-region
$0 / $0Free origin API; not FlashX
$0.08 / $0.48Stockholm · in-region

A free or retired origin model is not the same commercial service as a managed cloud offering. Rates above come from AWS pricing. The original Azure DeepSeek V4 Pro is $1.74/$3.48 Global, but its April snapshot does not match today’s Pro-0813. The original V4 Flash is $0.19/$0.51; neither rate is a quote for V4.1 Flash. Azure DeepSeek prices

Qwen3.8-Max and Flash have no verified same-model pay-per-token quote in native Bedrock or Foundry. Their origin prices are shown in the preceding section. Kimi K2.7 Code’s Foundry tariff remains unverified. GLM-5.3-Flash is not verified in native Bedrock or Foundry, but is available through the external Fireworks route below.

Current models through Fireworks on AWS Marketplace

A separate external API, billed by token. Prices below are Fireworks’ public Standard serverless rates. Marketplace-specific tariffs and negotiated terms were not independently verified.

Current modelOrigin APIFireworks public API
$1.32 / $3.96Peak
$1.32 / $3.96
$0.3 / $1.2Peak
$0.22 / $0.66
$2 / $6Singapore International
$2 / $6
$3 / $15
$3 / $15
$0.3 / $1.2
$0.3 / $1.2
$1.4 / $4.4
$1.4 / $4.4
$0.15 / $0.5
$0.15 / $0.5

AWS Marketplace offer · Fireworks residency. AWS billing does not establish that processing runs in an AWS region. Together AI also offers serverless consumption credits through Marketplace, including Qwen3.8-Flash; its listing says “Deployed on AWS: No”. EU confinement and the Marketplace payment terms require separate verification.

06

Working with a partner in Japan

KDDI and other AWS partners can help with procurement and delivery. Their commercial role is separate from the inference service.

CustomerAnritsu

Selects the model, cloud account and processing requirements.

Commercial / delivery partnerJapanese AWS partner

Can provide procurement, billing, setup and support.

Inference serviceBedrock or a named provider

Runs the selected model under that service’s deployment rules.

PartnerHow it operatesRelevance to this decision
AWS billing agency, account management, implementation and operations support. The cloudpack service involves KDDI and iret.
Commercial and operational support for AWS adoption and usage. KDDI service · iret AWS offering
AWS procurement and billing support, Japanese invoicing, technical support and cloud operations.
An alternative for AWS procurement, invoicing and technical support.
AWS billing/resale, implementation and operational services.
AWS consumption billing with implementation and operational services.

These are three documented examples, not an exhaustive list or a claim that all three run their own model inference. Purchasing in Japan does not make inference Japan-local. General AWS reseller discounts do not necessarily cover Bedrock or Marketplace model charges.

Sources and scope

Official sources are linked beside the relevant facts. These are the main catalogues and price references.

OpenAI model specificationsSol, Terra and Astra prices and context tiers
Anthropic pricing and cloud billingCurrent Sonnet rate, Foundry CCU and cache pricing
DeepSeek pricing and aliasesCurrent Pro/Flash service and peak-hour prices
Kimi model catalogueCurrent models and K2.5 discontinuation
Alibaba Model Studio pricingQwen3.8 regional rates and Coder pricing
MiniMax pay-as-you-go pricingM3 and older model rates
Z.ai pricingCurrent GLM models and legacy Flash distinctions
Amazon Bedrock regional compatibilityIn-region, geographic and Global routing
Amazon Bedrock pricingStandard on-demand token tariffs
Azure Direct models and versionsModel IDs and snapshots
Foundry deployment availabilityGlobal and Data Zone Standard tables
Foundry deployment typesEU/EFTA processing area
Foundry partner modelsClaude operator and deployment variants
Fireworks on FoundryPay-per-token versus provisioned offers
Foundry Fireworks pricingK3 and M3 tariffs; exact 5.3 price unresolved
Fireworks on AWS MarketplaceSeparate per-token purchasing route

Catalogue and pricing snapshot: 16 September 2026. Performance sources checked on 17 September 2026. Availability is based on public documentation; account access and capacity have not been tested.