Skip to main content
Anthropic’s Claude models bring advanced conversational AI capabilities to Microsoft Foundry, providing state-of-the-art language understanding and generation for intelligent applications. Claude models excel at complex reasoning, code generation, and multimodal tasks including image analysis. This article describes the available Claude models, how they’re hosted and billed, supported APIs, capabilities, and best practices. To deploy and call a Claude model, see Deploy and use Claude models in Microsoft Foundry.
Items marked preview in this article are currently in preview. This preview is provided without a service-level agreement, and Microsoft doesn’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.

How Claude models are hosted and billed

Microsoft Foundry offers Claude models in two versions:
  • Version 1: Hosted on Anthropic infrastructure; these models run on Anthropic’s infrastructure (outside of Azure).
  • Version 2: Hosted on Azure; these models run on Azure infrastructure end-to-end and are all Generally available (GA).
Not all models are available in both versions. A model’s lifecycle stage, such as Preview or Generally available, can differ between the two versions. For per-model availability and lifecycle status, see Available Claude models. To compare both hosting options across data residency, SLAs, support paths, compliance, and purchasing flow, see Compare Azure-hosted and Anthropic-hosted Claude models.
You access Claude models in Microsoft Foundry through Foundry Models from partners and community. Models from partners and community that Anthropic sells and operates are Non-Microsoft Products under the Product Terms.Claude models in Foundry require an Azure Marketplace subscription and bill through Claude Consumption Units (CCU). Ensure that you have the permissions required to subscribe to model offerings before you deploy. For pricing details, see Claude Consumption Units (CCU) billing in Microsoft Foundry.

Available Claude models

The following Claude models comply with the EU watermarking standard: claude-mythos-5-1, claude-fable-5-1, and claude-fable-5. While claude-fable-5 uses single-key watermarking, claude-mythos-5-1 and claude-fable-5-1 use double key (interwoven 2-key and C2PA) as follows:
  • Interwoven text watermarking (second key): a second key is added for interwoven watermark detection per EU standards. Watermarking happens server-side at generation time; there is no change to request or response shapes. A separate watermark detection API is in early access and isn’t part of the Foundry launch scope.
  • C2PA watermarking: C2PA content-provenance marking, handled server/client-side by Anthropic surfaces. No API shape change.
The following table compares model availability for both versions of Claude models in Foundry. For details on the features referenced in the table, see the Capabilities and advanced features section. For errors you might encounter when you deploy or call Claude models, see Deploy and use Claude models: Troubleshooting. 1 Claude Fable 5 and Claude Fable 5.1 apply extra input/output classifiers that might refuse requests if the content triggers dual-use safeguard policies. When a refusal happens, the request returns a successful (200) response with a refusal indicator stop_reason: "refusal" instead of model-generated content. You aren’t billed for input tokens that are refused. 2 Claude Mythos 5-1, Claude Mythos 5, and Claude Mythos Preview are only available as gated research preview. Access to the models is granted solely at Anthropic’s discretion and prioritized for defensive cybersecurity use cases. See the Claude Fable 5.1 & Claude Mythos 5.1 system card, Claude Mythos 5 system card, and Claude Mythos Preview system card for responsible use guidance. 3 Per-turn effort controls, Mid-conversation, and Token budgets are currently in Beta state.

API overview

The following table lists the APIs that you can use to interact with both the Hosted on Azure and Hosted on Anthropic infrastructure versions of Claude models in Foundry. Use the Anthropic SDKs and the following Claude APIs:
The Hosted on Anthropic infrastructure version of Claude models in Foundry supports more APIs than the ones listed in this table. You can see them on the Claude API docs: API overview page.
1You can call the Messages API from the anthropic Python package, the @anthropic-ai/foundry-sdk JavaScript package, or directly through REST. The deployment endpoint follows the shape https://<resource-name>.services.ai.azure.com/anthropic/v1/messages, and REST and JavaScript clients use the anthropic-version: 2023-06-01 header.

Capabilities and advanced features

Claude models in Foundry expose core capabilities for processing, analyzing, and generating content, and tools that let Claude interact with external systems, execute code, and perform automated tasks. Claude’s API surface is organized into five areas: The following sections and tables summarize capabilities available across the Hosted on Azure and Hosted on Anthropic infrastructure versions of Claude models in Foundry. Unless noted, a capability applies to both versions.
The Hosted on Anthropic infrastructure version of Claude models in Foundry supports more capabilities than the ones listed in these tables. You can see the full list of capabilities on Claude Platform Docs: Features overview.For more information about the available capabilities and advanced features for Claude models in Foundry, see the Microsoft Developer Blog.

Model capabilities

Ways to steer Claude and Claude’s direct outputs, including response format, reasoning depth, and input modalities.

Thinking and effort

The Thinking feature allows specific values for the thinking parameter type, depending on the model, as described in the following table. The adaptive type configures the adaptive thinking feature, allowing the model to decide whether to think, based on query complexity and effort level. For example, thinking={"type": "adaptive"}. 1 Thinking can be disabled only at effort high or below The Effort feature allows specific effort levels for each model, as described in the following table. The xhigh level produces the same result as max.

Tools

Let Claude take actions on the web or in your environment. This feature consists of built-in tools that Claude invokes through tool_use. The platform runs server-side tools, and you implement and execute client-side tools.

Tool infrastructure

Discover, orchestrate, and scale tool use.

Context management

Control and optimize Claude’s context window for long-running sessions.

Files and assets

Manage the documents and data you provide to Claude.

Agent support

Deployment types and regions

Claude models in Foundry are available for the following deployment types in specific Azure regions:
  • Global Standard: All Claude models (Hosted on Azure and Hosted on Anthropic infrastructure).
  • Data Zone Standard (US): Hosted on Azure versions of claude-opus-5-5, claude-opus-5, claude-opus-4-8, claude-sonnet-5, and claude-sonnet-5-5.
For the exact Azure regions where Claude models are available for deployment, see Region availability by deployment type.

Quotas and rate limits

Rate limits for Claude models vary by model, deployment type, hosting version, and Azure subscription type. Limits are measured in requests per minute (RPM), uncached input tokens per minute (ITPM), and output tokens per minute (OTPM). For current limits, shared quota behavior, and prompt cache accounting, see Claude model quotas and rate limits.

Responsible AI considerations

When using Claude models in Foundry, consider these responsible AI practices:

Best practices

Follow these best practices when working with Claude models in Foundry:

Prompt engineering

  • Clear instructions: Provide specific and detailed prompts.
  • Context management: Use the available context window effectively.
  • Role definitions: Use system messages to define the assistant’s role and behavior.
  • Structured prompts: Use consistent formatting for better results.

Cost optimization

To optimize your usage and avoid rate limiting:
  • Implement retry logic: Handle 429 responses with exponential backoff.
  • Batch requests: Combine multiple prompts when possible.
  • Monitor token usage: Track your token consumption and request patterns.
  • Use appropriate models: Use the most cost-effective model for your use case. See Available Claude models.