Premium frontier · OpenAI
GPT-6 Astra
For particularly demanding planning and development.
$10 input / $50 output
Per million origin API tokens. 2.5× Sol’s standard token rates. Model details
Amazon Bedrock & Microsoft Foundry
Model capabilities, cloud availability and token prices for planning and implementation in a coding agent.
Top-tier and everyday models for planning, implementation and review.
| Model developer | Top tier | Everyday workhorse | Coding performance |
|---|---|---|---|
OpenAIUnited States | GPT-5.6 SolTop-tier coding reference | GPT-5.6 TerraGeneral coding at lower cost | |
AnthropicUnited States | Claude Opus 5Complex implementation and review | Claude Sonnet 5Everyday agentic development | |
DeepSeekChina | V4-Pro-0813Current Pro snapshot · August 2026 | V4.1 FlashLatest Flash · September 2026 | Pro-0813SWE-Pro: — · TB 2.1: 87.9% V4.1 FlashSWE-Pro: — · TB 2.1: 90.6% |
Moonshot AIChina · Kimi | Kimi K3Flagship · 1M context | Kimi K2.7 CodeCoding specialist · 256K context | |
AlibabaChina · Qwen | Qwen3.8-MaxLatest snapshot: 0902 | Qwen3.8-FlashLower-cost 3.8 variant | Max · original releaseSWE-Pro: 67.7% · TB 2.1: 86.6%Not verified for 0902; SWE-Pro uses corrected tasks. FlashSWE-Pro: 62.5% · TB 2.1: — |
MiniMaxChina | MiniMax M3Latest flagship | MiniMax M3Also the economical choice | M3 · both rolesSWE-Pro: 59.0% · TB 2.1: 66.0% |
Z.aiChina · GLM | GLM-5.3Flagship · August 2026 | GLM-5.3-FlashLower-cost current variant |
SWE-Pro = SWE-bench Pro, repository-level software tasks. TB 2.1 = Terminal-Bench 2.1, tasks performed through a terminal. Higher percentages indicate more tasks solved within each evaluation. A dash means no result was verified for that exact model and benchmark; it is not a zero score.
Results are reported by the model developers under different agent setups, reasoning effort and evaluation conditions, so they are not a controlled ranking. Sonnet uses max effort for SWE-Pro and xhigh for TB 2.1; Qwen Max uses corrected SWE-Pro tasks. Terminal-Bench 3.0 is not comparable to 2.1. Each model label in the performance column links to its source.
Premium frontier · OpenAI
For particularly demanding planning and development.
$10 input / $50 output
Per million origin API tokens. 2.5× Sol’s standard token rates. Model details
Premium frontier · Anthropic
For complex, long-running coding work.
$10 input / $50 output
Per million origin API tokens. 2× Opus 5’s standard token rates. Model details
Astra and Fable are excluded from the savings baselines below. Their higher token rates can still be offset by fewer tokens, fewer retries or caching. Fable 5.1 cache reads cost $0.25 per million tokens.
Current origin models and the pay-per-token versions available through each cloud.
| Developer | Latest at the origin | Amazon Bedrock | Microsoft Foundry |
|---|---|---|---|
OpenAI | GPT-5.6 Sol / TerraCurrent practical pair | GPT-5.6 Sol / TerraSame model names; separate regional routing rulesCurrent family | GPT-5.6 Sol / TerraBoth: version 2026-07-09 · Azure DirectCurrent family |
Anthropic | Opus 5 / Sonnet 5Current practical pair | Opus 5 / Sonnet 5Same model names; EU-local options existCurrent family | Opus 5 / Sonnet 5GA. v1: Anthropic infrastructure; v2: Azure infrastructure. Anthropic operates both.Current family |
DeepSeek | V4-Pro-0813Current ProV4.1 FlashLatest Flash | V3.2Earlier than both current origin choicesEarlier release | V4 Pro2026-04-23; older than Pro-0813V4-Flash-07312026-07-31 · preview; earlier than V4.1Earlier snapshots |
Moonshot / Kimi | K3 / K2.7 CodeFlagship / coding workhorse | K2.5Earlier generation; retired from the origin APIEarlier release | K3 / K2.7 CodeFireworks partner · pay per tokenK2.7 CodeAlso Azure Direct · 2026-06-12 · previewCurrent names |
Alibaba / Qwen | 3.8-Max / 3.8-FlashMax snapshot: qwen3.8-max-0902 | Qwen3 Coder 480BLarge coding model; distinct from Qwen3.8Qwen3 Coder Next / 30BSmaller coding alternativesDifferent coding branch | No token offer verifiedQwen catalogue weights and provisioned offers do not establish a pay-per-token route.No PAYG verified |
MiniMax | M3Latest model; also the workhorse | M2.5Earlier than M3Earlier release | M3Fireworks partner · 512K cloud context; origin supports 1MCurrent name |
Z.ai / GLM | 5.3 / 5.3-FlashFlagship / workhorse | GLM 5 / GLM 4.7 FlashBoth earlier than the current pairEarlier releases | GLM-5.3Fireworks partner · pay per token5.3-Flash not listedGLM-5.2-Fast is a different alternative.Flagship only |
“Current” means the published model family or name matches, not that weights, context limits or serving behaviour are identical. Availability here is documented pay-per-token availability; account access and capacity were not tested.
| Model | Amazon Bedrock | Microsoft Foundry |
|---|---|---|
GPT-6 AstraGlobal routing | GPT-6 AstraVersion 2026-09-03 · Global Standard | |
Fable 5.1Global routing from European endpoints | Fable 5.1Version 1 · Anthropic-hosted preview |
Newer Chinese models can also be purchased through AWS Marketplace. Fireworks AI offers current DeepSeek V4.1 Flash, V4 Pro-0813, Qwen3.8-Max, Kimi K3, MiniMax M3 and GLM-5.3/Flash through its external per-token API. These are separate from the native Bedrock catalogue. See the model-by-model prices.
The location of an endpoint can differ from the location where a model processes the request.
The request enters the service in Europe. A Global deployment can still process it elsewhere.
In-region processing or a geographic deployment policy limits where inference runs. The exact area matters.
| Available models | European endpoint | Inference confined to the EU? |
|---|---|---|
Opus 5 / Sonnet 5 | Stockholm or Ireland | Yes, with in-region deploymentAvailable through in-region deployment and EU geographic profiles. Model routing |
Astra / Sol / Terra | Yes, Global routing | No on the documented European routesAn EU endpoint does not limit their Global inference. OpenAI routing |
Fable 5.1 | Yes, Global routing | No on the documented European routesFable routing |
DeepSeek V3.2 / Kimi K2.5 | Stockholm | |
MiniMax M2.5 / GLM 5 / GLM 4.7 Flash | Stockholm; more EU regions for some models | |
Qwen3 Coder 480B / 30B | Stockholm | Yes, in-regionCoder 480B · Coder 30B |
Qwen3 Coder Next | London is documented; EU entries conflict | Not verified for the EUPricing lists EU locations that the model card does not confirm. London is outside the EU. Model card |
| Available models | European endpoint | Inference confined to the EU? |
|---|---|---|
GPT-5.6 Sol / Terra | EU Data Zone is listed | EU/EFTA area, not EU27-onlyCurrent deployment policy includes EFTA. Terra’s Standard availability and tariff are not consistently documented. Deployment policy |
GPT-6 Astra | Yes, Global Standard | No EU Data Zone Standard listedRegional availability |
Fable 5.1 / Opus 5 / Sonnet 5 | Sweden Central · Global | No EU-confined offer verifiedThe European entry uses Global routing. Partner deployment table |
DeepSeek V4-FlashOriginal 2026-04-23 version | EU Data Zone Standard | EU/EFTA area, not EU27-onlyThis is an older Flash snapshot. Regional availability |
V4 Pro / Flash-0731 / K2.7 Code | Yes, Global Standard | No EU-confined route verifiedRefers to the Azure Direct models. Regional availability |
Fireworks K3 / K2.7 Code / M3 / GLM-5.3 | No European pay-per-token deployment listed | US Data Zone onlyThese Fireworks offers are excluded from the EU Data Boundary. Fireworks deployment table |
Qwen3.8 | No qualifying token offer verified | Not establishedA model-weight catalogue entry is not a serverless inference commitment. |
Native Bedrock offers several older Chinese models with EU-local processing. Foundry offers more recent Kimi, MiniMax and GLM models through Fireworks, with US processing.
Fireworks purchased separately through AWS Marketplace has no published EU-only guarantee for default serverless inference. Its published restrictions are unrestricted or US; other locations require a separate offer. Fireworks residency.
This section describes inference processing. It does not establish where all logs, abuse-monitoring records or other service data are retained. AWS routing · Azure deployment types
Origin API prices, with Opus 5 and Sol as top-tier references and Sonnet 5 and Terra as everyday references.
Indicative token costs. These are candidates for the same development roles. No matched evaluation on Anritsu’s tasks establishes that they can replace Opus 5 or Sol at equal quality.
All rates below are USD per million tokens at the model developer’s API. The illustrative bill uses 1 million input + 200,000 output tokens. It holds token volume fixed; it is not a measured cost per completed task.
| Model | Input / output | Illustrative token bill | Saving at equal token volume |
|---|---|---|---|
$5 / $25 | $10.00 | Reference model | |
$4 / $20 | $8.00 | Reference model | |
$1.32 / $3.96 | $2.11 | 79% less than Opus 574% less than Sol | |
$3 / $15 | $6.00 | 40% less than Opus 525% less than Sol | |
$2 / $6 | $3.20 | 68% less than Opus 560% less than Sol | |
$0.3 / $1.2 | $0.54 | 95% less than Opus 593% less than Sol | |
$1.4 / $4.4 | $2.28 | 77% less than Opus 572% less than Sol |
At the same token budget, these Chinese candidates cost 25–93% less than Sol and 40–95% less than Opus 5. The unweighted average saving across the five Chinese developers is 72% versus Opus 5. Kimi K3 sits at the expensive end; MiniMax M3 at the inexpensive end.
| Model | Input / output | Illustrative token bill | Saving at equal token volume |
|---|---|---|---|
$2 / $10 | $4.00 | Reference model | |
$2 / $12 | $4.40 | Reference model | |
$0.3 / $1.2 | $0.54 | 86% less than Sonnet 588% less than Terra | |
$0.95 / $4 | $1.75 | 56% less than Sonnet 560% less than Terra | |
$0.15 / $0.47 | $0.24 | 94% less than Sonnet 594% less than Terra | |
$0.3 / $1.2 | $0.54 | 86% less than Sonnet 588% less than Terra | |
$0.15 / $0.5 | $0.25 | 94% less than Sonnet 594% less than Terra |
For the everyday choices, the unweighted average saving across the five Chinese developers is 83% versus Sonnet 5, with individual savings of 56–94%. These are price comparisons within broad usage tiers, not measured equivalence in coding capability.
Basis: standard uncached rates, DeepSeek peak hours, Qwen Singapore International, and inputs within the lowest context tier. No batch discounts, taxes, tool charges or partner fees. Sol’s $4/$20 origin offer is stated through at least 21 November 2026. Sol terms
DeepSeek rates are half-price outside its published peak windows. Qwen Frankfurt Global is lower than Singapore International: Max $1.65/$4.951 and Flash $0.113/$0.382. A Frankfurt Global endpoint does not establish EU-only inference.
Large prompts can trigger higher tiers: OpenAI above 272K input tokens; MiniMax M3 above 512K. The table uses one million input tokens accumulated across requests within the base tier, not a single one-million-token prompt. Reasoning output, cache reads and writes, retries, and successful task completion all affect the actual bill. Different models can tokenize the same text differently.
Provider benchmarks use different harnesses, effort levels and task versions. They support a shortlist, but not a single cross-provider claim of equal coding capability. DeepSeek terms · Qwen pricing · MiniMax pricing
Rates depend on the model version, processing location and commercial offer.
Same published model names are compared below. USD per million input / output tokens. Standard Global rates unless a cell states US Data Zone; these are not EU-local quotes.
| Model | Origin API | Amazon Bedrock | Microsoft Foundry |
|---|---|---|---|
$4 / $20 | $4 / $20Global | $4 / $20Global · promotional | |
$2 / $12 | $2 / $12Global | $2 / $12Global | |
$5 / $25 | $5 / $25Global | $5 / $25Global · token-based CCU | |
$2 / $10 | $2 / $10Global | $2 / $10Global · token-based CCU | |
$3 / $15 | Not listed natively | $3.3 / $16.5Fireworks · US Data Zone · 10% higher | |
$0.3 / $1.2 | Not listed natively | $0.33 / $1.32Fireworks · US Data Zone · 10% higher | |
$1.4 / $4.4 | Not listed natively | Available via FireworksExact GLM-5.3 tariff not verified. Pricing |
Bedrock EU-local Opus 5 costs $5.50/$27.50; Sonnet 5 costs $2.20/$11, 10% above the Global rates. Astra and Fable 5.1 cost $10/$50 on all three Global routes. Foundry CCU is token-derived usage billing, not payment for hosted compute. AWS tariffs · Foundry Claude billing
| Exact model | Origin API today | Bedrock in the EU |
|---|---|---|
No longer served directlyOrigin release history | $0.74 / $2.22Stockholm · in-region | |
Discontinued on 31 AugustOrigin model catalogue | $0.72 / $3.6Stockholm · in-region | |
$0.861 / $3.441Frankfurt Global · ≤32K prompt | $0.45 / $1.8Stockholm · in-region | |
$0.3 / $1.2 | $0.36 / $1.44Stockholm · in-region | |
$1 / $3.2 | $1.2 / $3.84Stockholm · in-region | |
$0 / $0Free origin API; not FlashX | $0.08 / $0.48Stockholm · in-region |
A free or retired origin model is not the same commercial service as a managed cloud offering. Rates above come from AWS pricing. The original Azure DeepSeek V4 Pro is $1.74/$3.48 Global, but its April snapshot does not match today’s Pro-0813. The original V4 Flash is $0.19/$0.51; neither rate is a quote for V4.1 Flash. Azure DeepSeek prices
Qwen3.8-Max and Flash have no verified same-model pay-per-token quote in native Bedrock or Foundry. Their origin prices are shown in the preceding section. Kimi K2.7 Code’s Foundry tariff remains unverified. GLM-5.3-Flash is not verified in native Bedrock or Foundry, but is available through the external Fireworks route below.
A separate external API, billed by token. Prices below are Fireworks’ public Standard serverless rates. Marketplace-specific tariffs and negotiated terms were not independently verified.
| Current model | Origin API | Fireworks public API |
|---|---|---|
$1.32 / $3.96Peak | $1.32 / $3.96 | |
$0.3 / $1.2Peak | $0.22 / $0.66 | |
$2 / $6Singapore International | $2 / $6 | |
$3 / $15 | $3 / $15 | |
$0.3 / $1.2 | $0.3 / $1.2 | |
$1.4 / $4.4 | $1.4 / $4.4 | |
$0.15 / $0.5 | $0.15 / $0.5 |
AWS Marketplace offer · Fireworks residency. AWS billing does not establish that processing runs in an AWS region. Together AI also offers serverless consumption credits through Marketplace, including Qwen3.8-Flash; its listing says “Deployed on AWS: No”. EU confinement and the Marketplace payment terms require separate verification.
KDDI and other AWS partners can help with procurement and delivery. Their commercial role is separate from the inference service.
Selects the model, cloud account and processing requirements.
Can provide procurement, billing, setup and support.
Runs the selected model under that service’s deployment rules.
| Partner | How it operates | Relevance to this decision |
|---|---|---|
AWS billing agency, account management, implementation and operations support. The cloudpack service involves KDDI and iret. | Commercial and operational support for AWS adoption and usage. KDDI service · iret AWS offering | |
AWS procurement and billing support, Japanese invoicing, technical support and cloud operations. | An alternative for AWS procurement, invoicing and technical support. | |
AWS billing/resale, implementation and operational services. | AWS consumption billing with implementation and operational services. |
These are three documented examples, not an exhaustive list or a claim that all three run their own model inference. Purchasing in Japan does not make inference Japan-local. General AWS reseller discounts do not necessarily cover Bedrock or Marketplace model charges.
Official sources are linked beside the relevant facts. These are the main catalogues and price references.
Catalogue and pricing snapshot: 16 September 2026. Performance sources checked on 17 September 2026. Availability is based on public documentation; account access and capacity have not been tested.