The right AI provider is rarely "the one with the best model." It is the provider and model family that fits the job, data posture, integration surface, cost envelope, latency expectation, and operational controls of the system you are building.
This page maps major providers and their current model families as of September 15, 2026. Model IDs, availability, context limits, prices, and regional support change quickly; validate the linked first-party documentation before building or buying.
The selection lens
Start with the workload, then select a model tier. Do not begin with a brand preference.
At-a-glance reference
| Provider | Current families to know | Common architectural fit | Deployment posture |
|---|---|---|---|
| OpenAI | GPT-6 Astra; GPT-5.6 Sol, Terra, Luna; GPT Image 2.5 Sunburst and Flare; GPT-Live 1, realtime and transcription models; specialist research models | Frontier reasoning, coding, agentic work, multimodal and voice applications | Hosted API and platform services |
| Anthropic | Claude Fable 5.1; Claude Opus 5; Claude Sonnet 5; Claude Haiku 4.5; Claude Mythos 5.1 (invitation only) | Long-running agents, reasoning, coding, enterprise analysis | Hosted API and cloud marketplace access |
| Gemini 3 family (3.8 Flash current); Gemini 3.8 Live; Gemini 3.5 Transcribe; Gemini Omni Flash; Nano Banana image models; Veo; Lyria 3.5; Gemma | Multimodal applications, real-time voice, Google ecosystem, high-throughput work | Gemini API / Vertex AI; selected open models | |
| xAI | Grok 4.6 / 4.5; Grok 4.3 and 4.20 long-context snapshots; Grok Build; Grok Imagine; Grok Voice | Reasoning, coding, real-time and creative modalities | xAI API; selected models through Amazon Bedrock and Google Enterprise Agent Platform |
| Meta | Llama 4 Scout and Maverick | Custom hosting, controlled deployment, multimodal open-weight use cases | Open-weight under Meta license; multiple hosts |
| Mistral AI | Mistral Large 3; Mistral Medium 3.5; Mistral Small/Ministral families | Efficient enterprise inference, European deployment options, OCR and specialist work | API plus open-weight / self-host options |
| Cohere | Command A+; Command A (incl. Reasoning, Vision, Translate); North; Transcribe; Embed/Rerank/Parse; Aya | Enterprise RAG, tool use, coding, document parsing, multilingual and speech workloads | API and private/VPC/on-prem options |
| Amazon | Nova 2 family; legacy Nova v1 generation | AWS-native multimodal, document and media workflows, governed cloud delivery | Amazon Bedrock |
| Microsoft | Phi; MAI and healthcare models; Model Router and partner catalog | Microsoft/Azure enterprise platform, governed multi-provider choice, smaller-model scenarios | Microsoft Foundry / Azure |
| DeepSeek | DeepSeek V4.1 Flash; DeepSeek V4 Pro; DeepSeek R1 family | Cost-conscious reasoning/coding and multimodal evaluation, compatible API integrations | Hosted API; open-weight ecosystem for select releases |
| Alibaba Cloud / Qwen | Qwen 3.8-Max and 3.8-Flash; Qwen3.7-Plus; Qwen code, audio and multimodal families | Multilingual, coding, hybrid-thinking and open-model evaluation | Qwen API / Alibaba Model Studio; open-weight ecosystem |
Provider notes and model references
OpenAI
OpenAI now documents a generation above GPT-5.6. GPT-6 Astra (gpt-6-astra, released September 3, 2026) is listed as its most capable model, built for hard end-to-end work spanning reasoning, coding, computer use, research, and document creation. It carries a 1.05M-token context window with 128K max output at $10 / $50 per million input/output tokens, and input above 272K tokens bills at 2× the input rate and 1.5× the output rate — long-context work is a budget decision, not a free capability. Astra also adds agentic controls that change how an agent loop is written rather than only what it costs: async tool calling, mid-turn steering, and mid-conversation reasoning-effort adjustment.
The GPT-5.6 family remains current beneath it — Sol for demanding professional reasoning and coding, Terra for balanced capability and cost, and Luna for efficient, high-volume workloads — alongside dedicated image generation/editing and realtime/audio families.
The specialist and voice lines moved in the second week of September. On September 8, OpenAI released GPT-Image-2.5 Sunburst and GPT-Image-2.5 Flare — the most capable and the fast everyday image generation/editing options, adding xhigh and max quality settings at GPT Image 2 token rates — and made GPT-Rosalind Research generally available for life-sciences reasoning at $5 / $25 per million input/output tokens, with billing starting October 5, 2026 and access limited to approved organizations. On September 10, GPT-Live 1 reached general availability as a full-duplex voice model on the v1/live/sessions endpoint, priced at $0.05 per minute billed per second, with backend model and tool usage charged separately — voice pricing is per-minute plus per-token, so model it as two cost lines, not one. The same day added the Agents API in public beta (a managed Codex harness with durable sessions, progress streaming, and MCP/custom tool support, where OpenAI handles session orchestration, context compaction, and recovery) and project API key expiration with organization-enforced maximum key lifetimes. Prompt cache diagnostics went generally available on September 8 for GPT-5.6 and later on the Responses API, which makes cache-miss causes observable rather than inferred.
Platform-surface lifecycle continues to matter as much as model lifecycle. The legacy Assistants API reached its sunset on August 26, 2026 (its replacements are the Responses and Conversations APIs), and the removal of the Videos API and the Sora 2 model family is now nine days out — September 24, 2026, with no named replacement — a useful reminder that a provider can exit a modality you built on. On August 26, 2026, OpenAI also deprecated its transcription line: whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize shut down on February 26, 2027, with gpt-transcribe or gpt-live-transcribe as replacements. The gpt-3.5-turbo-instruct, babbage-002, davinci-002, and gpt-3.5-turbo-1106 snapshots retire on September 28, 2026.
Good starting point: Use a balanced model for normal product workflows; reserve the frontier tier for genuinely difficult reasoning, coding, or high-value decision support. A new top-of-range model is a reason to re-check routing, not to move everything up a tier — build model routing so a costly model is not used for simple extraction or classification.
Reference models: gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-image-2.5-sunburst, gpt-image-2.5-flare, gpt-live-1, gpt-transcribe and gpt-live-transcribe, GPT-Realtime family, gpt-rosalind-research where access applies. OpenAI model catalog · GPT-6 Astra model page · OpenAI deprecations · OpenAI changelog
Anthropic
Anthropic's documented current lineup is four models: Claude Fable 5.1 (released September 1, 2026) as the latest and highest-capability model, for demanding reasoning and long-horizon agentic work; Claude Opus 5 for complex agentic coding and enterprise work; Claude Sonnet 5 for the best combination of speed and intelligence; and Claude Haiku 4.5 as the fast, economical tier. Anthropic's own documented guidance is to start with Opus 5 for most workloads and move up to Fable 5.1 only when evals on Opus 5 at higher effort still fall short — a useful corrective to defaulting to the top of a provider's list. Fable 5 and Opus 4.8 are now among the legacy-but-available models. Claude Mythos 5.1 shares Fable 5.1's capabilities, specifications, and pricing but is offered by invitation only through Project Glasswing.
Several lifecycle details are worth carrying into a design review. Fable 5.1 makes three breaking changes against Fable 5: forced tool use (tool_choice of type any or tool) returns a 400 error, thinking blocks are bound to the model that produced them so no earlier model can read Fable 5.1's, and editing anything earlier in a conversation invalidates the thinking blocks after it — enforced for accounts created on or after August 31, 2026, and opt-in for older ones. Treating conversation history as append-only is now an integration requirement rather than a style preference. Migrating up from Opus 4.8 has its own hazard: Opus 5 turns adaptive thinking on by default, so a response can open with thinking blocks in code that assumed the first content block was text, and thinking tokens bill as output.
Two governance facts also carry forward. Fable 5.1 and Mythos 5.1 carry 30-day data retention, are not available under zero data retention absent express authorization from Anthropic, and include safety classifiers that can return a refusal as a successful HTTP 200 with stop_reason: "refusal" — so refusal handling and fallback belong in the integration, not in a runbook. Output from both also carries Anthropic's statistical text watermark, with C2PA Content Credentials on generated media retrieved through the Files API, which is a provenance fact worth recording in a content policy. Claude Opus 4.1 was retired on the Claude API on August 5, 2026, though Amazon Bedrock runs its own schedule and lists it to January 8, 2027; Bedrock's nearer date is Claude Sonnet 4, which reaches end of life there on October 14, 2026.
The model lineup has not changed since Fable 5.1's release, but two platform capabilities arrived in mid-September, and both affect how a long-running agent is built. On September 14, the Messages API added on-demand conversation compaction in beta (compact-2026-09-04): send a compaction parameter, get back a signed block summarizing earlier messages, and return that block in place of them on later requests, with recent turns preserved word-for-word and thinking still valid in those turns. That makes context budget a request-level control rather than something you hand-roll with your own summarizer. On September 10, Claude Managed Agents added an auto permission policy in which the server evaluates each agent or MCP tool call and runs it, denies it, or pauses for approval, reporting its decision in the event stream — an approval gate that lives in the platform rather than in your loop. Prompt cache reads on Fable 5.1 and Mythos 5.1 also cost 2.5% of base input price ($0.25 per MTok) rather than the usual 10%, which changes the arithmetic on cache-heavy agent designs at the top tier.
Good starting point: Consider Claude when long-context analysis, coding workflows, tool-using agents, and safety/process documentation are central to the use case. Evaluate outputs against the work your people actually perform; do not treat a capability claim as a production acceptance test.
Reference models: claude-fable-5-1, claude-opus-5, claude-sonnet-5, claude-haiku-4-5; claude-mythos-5-1 where access and terms apply. Claude models overview · Claude model selection · What's new in Claude Fable 5.1 · Anthropic deprecations · Claude API release notes
Google's portfolio spans Gemini for text, multimodal intelligence, and image generation (the Nano Banana models), Veo for video, Lyria for music, and Gemma for open models. The Gemini API's model guide now leads with Gemini 3.8 Flash (gemini-3.8-flash), generally available on September 2, 2026 and described as Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Gemini 3.7 Flash, 3.6 Flash, 3.5 Flash, and the Flash-Lite variants all remain stable and listed, so this is a new option at the top of the Flash line rather than a forced migration.
Real-time voice is the newest addition. On September 15, Gemini 3.8 Live (gemini-3.8-live) and Gemini 3.8 Live Extended Thinking (gemini-3.8-live-extended-thinking) became generally available as audio-to-audio models for real-time voice applications, the second supporting background reasoning during a conversation. A voice model that thinks mid-turn is a latency and interaction-design decision as much as a quality one; evaluate it against your own turn-taking expectations rather than a demo.
Two other changes landed earlier in the month. On September 1, agentic video understanding launched for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite across the Interactions and GenerateContent APIs, using up to 88% fewer tokens on long-form content than static processing — a cost lever for anyone already doing video analysis. On September 3, Lyria 3.5 (lyria-3.5) was released as generally available for full-length song generation with improved musical coherence and structural control; it is now Google's flagship music model. Elsewhere in generative media the labels differ: the Nano Banana image models (gemini-3.1-flash-image, gemini-3.1-flash-lite-image, and gemini-3-pro-image) are GA, Imagen 4 is deprecated, and Veo 3.1 remains preview-only — worth checking per model rather than per portfolio.
Late August moved two specialist lines from preview to general availability. On August 26, Gemini 3.5 Transcribe and Gemini 3.5 Transcribe Live became generally available as dedicated speech-to-text models, covering 85+ languages with speaker diarization, word-level timestamps, and custom vocabulary biasing — the streaming variant runs over the Live API. On August 27, Gemini Omni Flash reached GA as gemini-omni-1.1-flash, adding video extension, image-to-video interpolation, and explicit resolution control; the gemini-omni-flash-preview endpoint it replaces is scheduled for deprecation on September 30, 2026. If you prototyped on that preview endpoint, the migration window is now short.
Good starting point: Gemini is a natural evaluation candidate for multimodal workflows — especially when a system must interpret combinations of text, image, video, or audio. Keep the API surface and the enterprise/Vertex deployment surface distinct when comparing governance and data controls.
Reference models: gemini-3.8-flash, Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash and Flash-Lite; gemini-3.8-live and gemini-3.8-live-extended-thinking; gemini-3.5-transcribe and gemini-3.5-transcribe-live; gemini-omni-1.1-flash; lyria-3.5; gemini-3.1-flash-image and gemini-3-pro-image; Veo 3.1 (preview); Gemma. Gemini model guide · Gemini API release notes · Google AI developer platform
xAI
xAI's API exposes the Grok family across text, reasoning, coding, image/video generation, and voice. Its current API listing recommends Grok 4.6 as the default for everything including code, with Grok 4.5 alongside it; both carry a 500K context window and price input above 200K tokens at double the base rate. The listing also keeps Grok 4.3 and the Grok 4.20 reasoning, non-reasoning, and multi-agent snapshots, which have 1M-token context windows at lower per-token prices — so at xAI the longest context currently sits below the flagship, not at it. Grok Build is a coding model intended for agentic development tasks. Grok 4.6 became generally available through Amazon Bedrock on August 19, 2026, and available through Google Enterprise Agent Platform on August 21, expanding deployment choices beyond direct xAI access.
One deprecation is worth copying into a design review because it does not fail loudly: grok-imagine-image-quality retires on November 2, 2026, and from that date requests to it are routed to grok-imagine-image-2.0 with quality set to low at lower pricing, keeping the same request and response format. Code that keeps working while quietly producing different output is harder to catch than code that starts returning errors.
Good starting point: Evaluate Grok when your product benefits from a broad multimodal API or when coding-agent capabilities are central. As with any emerging platform, confirm data handling, regional access, retention, and model lifecycle before making it a core operational dependency.
Reference models: grok-4.6, grok-4.5, grok-4.3, grok-4.20-0309-reasoning and grok-4.20-multi-agent-0309, grok-build-0.1, grok-imagine-image-2.0, Grok Imagine video, Grok Voice. xAI model documentation · xAI release notes · Grok 4.6 on Amazon Bedrock · Grok 4.6 on Google Enterprise Agent Platform
Meta
Meta's Llama family is important because it provides a widely deployed open-weight path. Llama 4 Scout and Llama 4 Maverick are natively multimodal models for text and image understanding. This is a different procurement and operations decision from using a hosted frontier API: the organization takes on more responsibility for serving, observability, updates, safety controls, and license review.
Good starting point: Consider Llama when self-hosting, data locality, customization, offline/edge constraints, or provider flexibility are material requirements. "Open weight" is not synonymous with no obligations — review the specific license, model card, hosting configuration, and security posture.
Reference models: Llama 4 Scout, Llama 4 Maverick. Meta Llama 4 overview · Meta AI developer documentation
Mistral AI
Mistral offers both commercial API models and open-weight options, making it relevant when model portability and deployment choice matter. Mistral Medium 3.5 is its frontier-class multimodal model for agentic and coding work, while Mistral Large 3, Mistral Small 4, and the Ministral 3 sizes ship under Apache 2.0 — an unusually permissive license position for models at that capability level. The specialist catalog is where a lot of the practical value sits: OCR 4.1 for document extraction, Codestral for code, Voxtral TTS and Voxtral Mini Transcribe 2 for speech, and Shieldstral for moderation.
Good starting point: Put Mistral on the shortlist when you need a strong general model with options beyond a single hosted API, or when OCR/document processing is part of the application. Check the current model card rather than assuming a prior model remains active.
Reference models: mistral-large-2512 (Mistral Large 3), Mistral Medium 3.5, Mistral Small 4, Ministral 3 (14B / 8B / 3B), Mistral OCR 4.1, Codestral, Voxtral TTS and Voxtral Mini Transcribe 2, Shieldstral. Mistral model overview · Mistral Large 3 model card
Cohere
Cohere is oriented toward enterprise language workloads, retrieval-augmented generation (RAG), tool use, multilingual work, and flexible deployment. The current catalog includes Command A+, Command A Reasoning, Command A Translate, Command A Vision, and Command R7B; the North family is now North Mini Code for agentic coding and North Small Translate for translation. Cohere also maintains Transcribe (including an Arabic-optimized model), Embed, Rerank v4.0 in Pro and Fast variants, a Parse model for document extraction, and the Aya research and Tiny Aya multilingual families.
Good starting point: Cohere deserves careful evaluation for knowledge-grounded enterprise assistants, tool-using workflows, multilingual support, and systems that need a more controlled deployment story. Citation behavior and retrieval quality should be measured with your own knowledge corpus.
Reference models: command-a-plus-05-2026, command-a-reasoning-08-2025, command-a-translate-08-2025, command-a-vision-07-2025, command-r7b-12-2024, north-mini-code-1-0, north-small-translate-1-0, cohere-transcribe-03-2026, embed-v4.0, rerank-v4.0-pro and rerank-v4.0-fast, parse-v5.0, Aya Expanse and Tiny Aya. Cohere model list · Cohere model overview · Cohere release notes
Amazon
Amazon's Nova family is an AWS-native option accessed through Amazon Bedrock. Nova spans multimodal understanding, creative generation, and speech, with Nova 2 as the current generation. The earlier Nova v1 generation is now past the planning stage: Nova Premier v1 and Nova Sonic v1 reached end of life on September 14, 2026, and Nova Canvas v1 and Nova Reel v1/v1.1 follow on September 30, 2026. Bedrock publishes its own dates, which can differ from the model provider's.
Bedrock also split its lifecycle policy in two, and which half applies depends on when a model launched. Models launched on or after September 7, 2026 follow a new policy: every model card carries an "EOL no sooner than" date plus a declared Legacy notice period of either six months or 45 days, and the EOL date is added to the card when the model enters Legacy. Models launched before that date stay under the previous policy, which adds a "public extended access" phase — after at least three months in Legacy, active users can keep calling the model until EOL, but at higher pricing set by the model provider. Under both policies, new customers cannot adopt a Legacy model and existing customers can lose access after 15 days of inactivity, so an occasionally used fallback model is not a safe fallback. The practical consequence is that a 45-day notice window is a real possibility now; read the model card, not the general policy page, before you depend on a model.
Good starting point: Consider Nova when your data, identity, audit, networking, and deployment workflows already live in AWS. Bedrock can also be evaluated as a model-access and governance layer rather than as a commitment to a single provider.
Reference models: Nova 2 family. Nova Premier v1 and Nova Sonic v1 are retired; treat Nova Canvas v1 and Nova Reel v1/v1.1 as migrations to finish this month, not new-build options. Amazon Nova overview · Amazon Bedrock model lifecycle · Bedrock model lifecycle (legacy policy)
Microsoft
Microsoft's differentiator is the enterprise platform around models: Microsoft Foundry provides a catalog, build, evaluation, deployment, and governance surface for Microsoft, OpenAI, Anthropic, Meta, Mistral, Cohere, DeepSeek, xAI, and other models. Microsoft's own families include Phi, MAI, and healthcare-oriented models. Foundry Model Router expanded to 28 regions in August 2026, adding GPT-5.6 Sol/Terra/Luna and Claude Opus 4.8 while removing several retired model versions from its routing pool; in September 2026 it reached 32 regions (adding Canada Central, North Europe, Norway East, and UAE North) and added preview per-request routing metadata reporting the routing mode, ordered model attempts, and any fallback that occurred. That observability matters: a router is only governable if you can see which model actually served a request. September also added preview session affinity for Chat Completions — supply an opaque session ID and the router will try to keep related conversation turns on the same eligible model, reporting whether the association was initialized, retained, or switched. Routing that changes model mid-conversation is a consistency problem as much as a cost one, so this is the control to reach for before you pin a single model and give up routing entirely.
Good starting point: Consider Foundry when Azure identity, networking, compliance controls, and multi-provider governance are more important than direct-to-provider access. Microsoft is often the platform choice, not a single-model answer — confirm the exact model, region, deployment type, and commercial terms.
Reference models: Phi-4 and current Phi family; MAI family; healthcare models; Azure OpenAI offerings; partner models in the catalog; Model Router where policy-based multi-model routing fits. Microsoft Foundry documentation · Model Router updates
DeepSeek
DeepSeek is relevant for its cost-conscious model ecosystem and API compatibility. On September 10, 2026, it released DeepSeek V4.1 Flash, described as the smallest model in a new architecture family, with native multimodal visual understanding, reached through the deepseek-flash ID and accompanied by API price reductions. Two lifecycle details matter more than the benchmark scores. First, the older deepseek-v4-flash and deepseek-v4-flash-vision-exp names now route to V4.1 Flash for compatibility, so a pinned model string here does not pin a model version — if you need reproducibility, version your evaluations rather than trusting the ID. That also resolves the experimental vision model referenced in earlier revisions of this page: it is a redirect now, not a separate model. Second, DeepSeek extended API service for DeepSeek V4 Pro past its previously announced September 14, 2026 cutoff, keeping current billing, which is a reprieve rather than a commitment. The V3 and R1 families remain useful references for reasoning and open-weight ecosystem evaluation.
Good starting point: Use DeepSeek as an evaluation option for coding, reasoning, and compatible API integrations, but complete a deliberate legal, data-residency, security, and vendor-risk review before placing sensitive operational data in the path.
Reference models: deepseek-flash (DeepSeek V4.1 Flash), deepseek-v4-pro, DeepSeek R1 family. DeepSeek API updates · DeepSeek models and pricing · DeepSeek model listing
Alibaba Cloud / Qwen
Alibaba Cloud's Qwen portfolio includes hosted and open-model paths covering text, code, vision, audio, video, and research-oriented workloads. The Qwen3 family supports selectable thinking and non-thinking modes; Qwen 3.8-Max is described by Qwen as its most capable family model for coding and collaborative work, with Qwen 3.8-Flash as the economical tier. Model Studio's featured list now pairs those with Qwen3.7-Plus for vision-language and GUI-agent work, alongside the qwen-audio-3.0 speech models (TTS, streaming and file transcription, and speech-to-speech), the qwen3.5-omni-plus omni-modal models, the qwen-image-3.0-pro and wan image/video generation families, and a set of embedding and reranking models for retrieval work.
Good starting point: Evaluate Qwen for multilingual applications, coding workloads, hybrid reasoning modes, or an open-model strategy. Treat deployment geography, model access method, and policy requirements as first-order architecture decisions.
Reference models: qwen3.8-max, qwen3.8-flash, qwen3.7-plus, qwen3.5-omni-plus, Qwen code family, qwen-audio-3.0 speech models, qwen-image-3.0-pro, wan3.0-video, text-embedding-v4 and qwen3-rerank. Qwen API platform · Qwen 3.8-Max release · Alibaba Model Studio model guide
How to evaluate a provider responsibly
Do not select a model from a marketing claim or a broad public benchmark alone. Use a small, representative evaluation pack drawn from the actual work you hope to improve:
- Task quality: Does it solve the real task accurately and consistently?
- Grounding: Does it cite or point to the correct internal evidence when required?
- Failure behavior: Does it surface uncertainty, refuse unsafe actions, and avoid inventing facts?
- Latency and cost: Does it meet the interaction and operating-cost target at realistic volume?
- Integration: Does it support structured outputs, tool use, modalities, and authentication patterns you need?
- Data and governance: Are retention, residency, audit, access, and vendor terms acceptable?
- Operational control: Can you trace decisions, monitor quality, route exceptions, and replace the provider if needed?
For operational documentation, compliance, hiring, safety, finance, or other consequential workflows, keep a qualified human responsible for the decision. AI may prepare evidence, flag gaps, or draft recommendations; it should not silently become the accountability layer.
Quick selection patterns
| Need | Sensible first evaluation set |
|---|---|
| High-stakes reasoning/coding with a hosted API | OpenAI, Anthropic, Google, xAI |
| High-volume classification, extraction, or summarization | Economical tiers from OpenAI, Google, Amazon, Cohere, Mistral, DeepSeek, Qwen |
| Document/image/video understanding | Google Gemini, OpenAI, Amazon Nova, Mistral, Cohere Vision, Meta Llama 4 |
| Real-time voice and transcription | OpenAI GPT-Live and GPT-Transcribe, Google Gemini 3.8 Live and 3.5 Transcribe, xAI Grok Voice, Cohere Transcribe, Mistral Voxtral, Qwen audio |
| Enterprise RAG and tool use | Cohere, Anthropic, OpenAI, Google, Amazon Bedrock, Microsoft Foundry |
| Self-hosting or deployment control | Meta Llama, Mistral, Qwen, DeepSeek, Microsoft Phi |
| Azure-centered governance | Microsoft Foundry and the models available in its catalog |
| AWS-centered governance | Amazon Bedrock and its available provider catalog |
| Multilingual focus | Google, Cohere, Mistral, Qwen, Meta Llama, OpenAI, Anthropic — validated on your languages and tasks |
The durable rule
Model selection is an architecture decision with a short shelf life. Keep a small abstraction boundary around the provider, version your prompts and evaluations, record the rationale in an architecture decision record, and revisit the choice as capabilities and terms evolve.
The strongest AI systems are not built around allegiance to a model brand. They are built around a clear task, reliable evidence, disciplined evaluation, and accountable operational ownership.
Change log
September 15, 2026
- Added OpenAI's September 8 and September 10 releases: GPT-Image-2.5 Sunburst and Flare with the new
xhighandmaxquality settings, GPT-Rosalind Research GA for life sciences (with billing starting October 5, 2026 and access limited to approved organizations), GPT-Live 1 GA for full-duplex voice at $0.05 per minute plus separate backend model and tool charges, the Agents API public beta, prompt cache diagnostics GA, and organization-enforced API key expiration. Updated the reference list from GPT Image 2 to the current image, voice, and transcription IDs, and flagged the Videos API / Sora 2 removal as nine days out rather than merely scheduled. - Recorded no change to Anthropic's four-model lineup, and added the two platform changes that do affect agent design: on-demand Messages API compaction (September 14, beta
compact-2026-09-04), which makes context budget a request-level control, and the Managed Agentsautopermission policy (September 10), which moves the tool-call approval gate into the platform. Noted the 2.5% prompt-cache read rate on Fable 5.1 and Mythos 5.1, and Amazon Bedrock's nearer Claude date — Sonnet 4 end of life on October 14, 2026. - Added Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking GA (September 15) for real-time audio-to-audio voice work, and corrected Lyria 3.5 from public preview to generally available. Noted that generative-media lifecycle labels differ per model — Nano Banana image models GA, Imagen 4 deprecated, Veo 3.1 still preview.
- Corrected the xAI section to match its current model listing: Grok 4.6 is the recommended default at 500K context, while Grok 4.3 and the Grok 4.20 snapshots carry 1M-token context at lower prices, so the longest context sits below the flagship. Added the
grok-imagine-image-qualityretirement on November 2, 2026, where requests are silently rerouted togrok-imagine-image-2.0atquality: lowinstead of failing. - Rewrote the Amazon section around Bedrock's split lifecycle policy: models launched on or after September 7, 2026 carry an "EOL no sooner than" date with a six-month or 45-day Legacy notice period, while earlier models keep the public extended access phase at provider-set higher pricing. Moved Nova Premier v1 and Nova Sonic v1 to retired (September 14, 2026) and left Nova Canvas v1 and Nova Reel v1/v1.1 as September 30 migrations. Added the 15-day inactivity rule, because it means a rarely exercised fallback model is not a real fallback.
- Updated DeepSeek to V4.1 Flash (September 10, 2026), reached via
deepseek-flashwith price reductions, and recorded that the olderdeepseek-v4-flashanddeepseek-v4-flash-vision-expnames now route to it — a pinned ID no longer pins a version. Also noted DeepSeek V4 Pro service continuing past its September 14, 2026 cutoff. - Added Microsoft Foundry Model Router's preview session affinity for Chat Completions, which addresses mid-conversation model switching without giving up routing.
- Refreshed the Cohere references against the current model list, replacing North Micro Vision — no longer listed — with North Small Translate, and adding Command A Reasoning, Parse, Rerank v4.0 Pro/Fast, the Arabic transcription model, and Tiny Aya. Refreshed Mistral with its Apache 2.0 licensing position and named specialist models (OCR 4.1, Codestral, Voxtral, Shieldstral), and Qwen with its current image, video, embedding, and reranking IDs.
- Added a real-time voice and transcription row to the quick selection patterns, since three providers shipped voice or transcription changes in this period.
- Confirmed no material change this period for Meta; that section stands as written.
September 7, 2026
- Added OpenAI's GPT-6 Astra (
gpt-6-astra, released September 3, 2026) as a generation above GPT-5.6, with its context window, pricing, the 272K-token long-context surcharge, and the async tool calling / mid-turn steering / mid-conversation effort controls that change how an agent loop is written. GPT-5.6 Sol, Terra, and Luna are now described as current tiers beneath it rather than as the frontier family. - Added OpenAI's August 26, 2026 transcription deprecation (
whisper-1,gpt-4o-transcribe,gpt-4o-mini-transcribe,gpt-4o-transcribe-diarize, shutting down February 26, 2027) and reconfirmed that the Videos API and Sora 2 removal is still set for September 24, 2026. - Updated Anthropic to Claude Fable 5.1 (released September 1, 2026) as the current top model, with Fable 5 moving to legacy-but-available and Claude Mythos 5.1 replacing Mythos 5 as the invitation-only Project Glasswing offering. Recorded Anthropic's documented guidance to start with Opus 5 for most workloads, because it argues against defaulting to the top of a provider's list.
- Replaced the Anthropic integration notes with Fable 5.1's three breaking changes — forced tool use returning a 400, thinking blocks bound to the producing model, and history edits invalidating later thinking blocks (enforced for accounts created on or after August 31, 2026) — and added the text watermark and C2PA Content Credentials as a provenance fact. Kept the Opus 4.8 → Opus 5 adaptive-thinking hazard and the 30-day retention, no-zero-data-retention, and refusal-as-HTTP-200 terms, which now apply to Fable 5.1 and Mythos 5.1.
- Added Google's Gemini 3.8 Flash GA (September 2, 2026), agentic video understanding for Gemini 3.7/3.6 Flash and 3.5 Flash-Lite (September 1), and Lyria 3.5 in public preview (September 3), while noting that the earlier Flash and Flash-Lite models remain stable — a new top option, not a forced migration.
- Updated Microsoft Foundry Model Router to 32 regions and added its preview per-request routing metadata, since routing is only governable when you can see which model served a request.
- Sharpened the Amazon Nova v1 end-of-life language from future planning to immediate action, with Nova Premier v1 and Nova Sonic v1 retiring September 14, 2026.
- Confirmed no material change this period for xAI, Meta, Mistral, Cohere, and DeepSeek, and no change to the featured Qwen lineup; those sections stand as written.
August 31, 2026
- Corrected the Anthropic section to the four-model lineup Anthropic currently documents — Claude Fable 5, Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5 — after verifying that Claude Opus 5 superseded Opus 4.8 at the Opus tier and that Fable 5 is a widely released model rather than a controlled-access one. Mythos 5 is now identified as the only access-gated offering.
- Added the Claude Opus 5 thinking-on-by-default breaking change and the Fable 5 refusal/data-retention terms, because both change integration and governance work rather than just model choice.
- Noted OpenAI's August 26, 2026 Assistants API sunset and the September 24, 2026 removal of the Videos API and Sora 2 family, which is a provider exiting a modality and not only a model swap.
- Added Google's August 26 GA of Gemini 3.5 Transcribe and Transcribe Live and the August 27 GA of Gemini Omni Flash, with the September 30, 2026 deprecation of the preview omni endpoint.
- Refreshed the Qwen references to the Model Studio featured list, replacing Qwen3.6 with Qwen3.7-Plus and adding Qwen 3.8-Flash and the audio families.
- Confirmed no material change this period for xAI, Meta, Mistral, Cohere, Amazon, Microsoft, and DeepSeek; those sections stand as written.
August 24, 2026
- Updated OpenAI GPT-5.6 Sol pricing/lifecycle guidance and corrected the primary model-catalog link.
- Reframed Anthropic around its broadly available Sonnet 5, Opus 4.8, and Haiku 4.5 lineup; retained Fable/Mythos as specialized-access references.
- Added Grok 4.6 availability through Amazon Bedrock and Google Enterprise Agent Platform.
- Expanded Cohere references to include North, Transcribe, Embed, and Rerank families.
- Marked affected Amazon Nova v1 models as legacy with September 2026 end-of-life dates.
- Updated Microsoft Foundry Model Router coverage and current routing pool.
- Added DeepSeek V4 Flash Vision Experimental with an explicit production-lifecycle caution.

