Skip to main content
The healthcare AI models marked (preview) in this article are currently in limited preview. These models are intended and provided as-is for research and model development exploration. The healthcare AI models are not designed or intended to be deployed in clinical settings as-is. They are not intended for use in the diagnosis or treatment of any health or medical condition, and the individual models’ performances for such purposes have not been established.You bear sole responsibility and liability for any use of the healthcare AI models, including verification of outputs and incorporation into any product or service intended for a medical purpose or to inform clinical decision-making, compliance with applicable healthcare laws and regulations, and obtaining any necessary clearances or approvals.
Fine-tuning adapts a premium healthcare AI model to your institution’s data distribution and reporting conventions while the model’s fundamental task stays the same. Its output aligns more closely to your label vocabulary and institutional style. To fine-tune a model, you prepare and upload a model-specific training file, create a fine-tuning job, and deploy the result. A job can remain pending until training capacity is available, moves to running, and then reaches a terminal status.
Premium healthcare models use the OpenAI Python client’s Files and fine-tuning interfaces after you obtain the client through the Foundry project SDK. Customize a model with fine-tuning covers fine-tuning OpenAI models with the Foundry SDK; this article covers premium healthcare models. Model names, schemas, supported training types, hyperparameters, and deployment details differ for premium healthcare models. Use the fine-tuning page for your selected healthcare model for model-specific information.

Prerequisites

Foundry stores uploaded training and validation data as needed to provide fine-tuning. Your training data and fine-tuned models aren’t available to other customers or model providers. Your training data isn’t used to train AI foundation models without your permission or instruction. For details, see Data, privacy, and security for Foundry Models sold by Azure.

Overview of the fine-tuning process

This process is common to all premium healthcare models. Each model has unique training data schemas, input and preprocessing requirements, supported hyperparameters, and evaluation guidance. Use the individual model page for model-specific information. Current model-specific guides:
  1. Verify regional support for inference deployments and fine-tuning jobs. Check training capacity and deployment quota separately.
  2. Prepare training data in the model-specific format.
  3. Upload the training file and, optionally, a validation file.
  4. Create the job using model-specific hyperparameter settings.
  5. The job remains pending until capacity is available, moves to running, and ends in succeeded, failed, or cancelled status.
  6. Deploy the fine-tuned model as a distinct deployment.
  7. Evaluate whether the fine-tuned model meets your application requirements.

Capacity, quota, and regional availability

Check fine-tuning availability and deployment quota

Each premium healthcare model has its own quota for base model and fine-tuned model deployments. Models don’t share quota with each other, and base and fine-tuned deployments don’t share quota. Each quota query is scoped to a subscription and region. To determine whether you can fine-tune a particular premium healthcare model in a region, check for its model-specific AIServices.GlobalStandard.<model-name>-finetune regional quota row. If the row is absent for your subscription and region, you can’t fine-tune the model there. One CxrReportGen Premium base model test observed usage accounted subscription-wide across regions. If base model quota usage doesn’t reconcile with deployments in the queried region, check that model’s deployments across the subscription. All premium healthcare models use the GlobalStandard deployment SKU. To query the model-specific quota rows for a subscription and region:
The quota row without -finetune governs base-model deployments. The row with -finetune governs live deployments of fine-tuned models. Its numeric limit is deployment capacity, not training capacity. This command returns the quota limit, not current usage or remaining capacity.
The -finetune row identifies regional fine-tuning availability, but its numeric value doesn’t determine how much you can train. Training capacity is configured separately on the AIServices account, and training jobs don’t consume this deployment quota.

Regional availability

The following regions are officially supported for all premium healthcare models. The combined fine-tuning column covers both fine-tuning jobs and fine-tuned model deployments. Experimental fine-tuning support is available in additional regions. Contact your Microsoft account team for more information.

Prepare your data

Create a required training file and an optional validation file, each using the JSONL messages envelope used by the fine-tuning service.

Anatomy of a single record

Each physical line in a JSONL file is one record with a top-level messages array. The exact roles, order, image count, text, and output shape are model-specific. The following JSONC schematic shows the shared envelope. It isn’t a complete schema. The object is expanded across lines for readability; in a JSONL file, store it on one physical line.
For the exact record schemas and examples, see CxrReportGen Premium training data format and MedImageInsight Premium training data format. Before uploading, check these shared requirements:
  • JSONL format — one record per line.
  • UTF-8 encoding.
  • Direct Files API upload and import operations require files smaller than 512 MB.
The 512 MB limit applies to the upload method, not the training-data format. For files at or above 512 MB and through 9 GB, use the multipart Uploads API.
For model-specific image requirements, see CxrReportGen Premium image preprocessing and requirements and MedImageInsight Premium image preprocessing and requirements.
Follow the CxrReportGen Premium training data requirements or MedImageInsight Premium training data requirements for the model you selected. You’re responsible for data suitability, de-identification, retention, access control, and compliance. This article doesn’t cover healthcare data governance.

Upload the training file

Upload your training file before creating a job. With the SDK, use the Foundry project client. For direct cURL requests, use the AIServices resource-host Files endpoint. A successful upload returns a file_id; wait until the file reaches processed status before creating a fine-tuning job.
The upload occurs in the fine-tuning job wizard:
  1. Start a fine-tuning job as described in Create the fine-tuning job.
  2. In Datasets, under Training data source, select an existing dataset or use Upload or drag and drop.
  3. Optionally add a Validation data source, and then select Next.
For request parameters, response fields, and other file operations, see the Upload file REST API reference. To import a file from Azure Blob Storage or a web location instead of uploading it directly, see the Import file REST API reference. For files that are 512 MB or larger, use the multipart Uploads API instead.

Create the fine-tuning job

Create a fine-tuning job after the training file reaches processed status. You can create the job in the Foundry portal or programmatically. All premium healthcare models configure n_epochs, batch_size, and learning_rate_multiplier under method.supervised.hyperparameters. Accepted ranges and service defaults vary by model. See the CxrReportGen Premium hyperparameters or MedImageInsight Premium hyperparameters for the model you selected.
The request body must include "trainingType": "globalStandard". Omitting trainingType causes the service to default to Standard, which is rejected with 400 invalidPayload.
  1. In your Foundry project, select Build > Fine-tune, and then select Start a new fine-tuning job.
  2. In Basic details, select a supported Customization method, the Model, and a supported Training type. Select Next.
  3. In Datasets, attach the required Training data source by selecting an existing dataset or Upload or drag and drop. Optionally, attach a Validation data source. Select Next.
  4. In Optional settings, enter a Display name and, optionally, a Seed. Under Hyperparameter tuning, review Batch size, Number of epochs, and Learning rate multiplier, and select Default or Custom for each value. Select Submit.
  5. After submission, the job appears in the Fine-tune jobs table.
For shared fine-tuning operations, request and response fields, job events, and cancellation, see the Fine-tuning REST API reference.

Monitor the fine-tuning job

In the Foundry portal, select Build > Fine-tune in your project. Use the jobs table to monitor a job, and open the job details page for more information.
A job stays pending while it waits for training capacity, moves to running when capacity is available, and ends in succeeded, failed, or cancelled. If it succeeds, keep the fine_tuned_model value for deployment. Check the job events for warnings or errors. For complete job statuses, response fields, events, and cancellation, see the Fine-tuning REST API reference.

Deploy the fine-tuned model

Deploy the fine-tuned model before you run inference against it. Use a deployment name different from your base model deployment. Before running the deployment command, verify these shared requirements:
  • Deployment name: different from any existing deployment in that account.
  • Model identifier (--model-name): the fine_tuned_model value from the succeeded job, in the format <Model>.ft-<jobhash>-<suffix>.
  • Version: always 1 for fine-tuned models.
  • Format: Microsoft.
  • SKU: GlobalStandard.
  • Capacity: a value within the available fine-tuned-deployment quota for your model and region, such as 1000.
  1. Open the completed job details page, and then select Deploy the fine-tuned model.
  2. Enter a Deployment name.
  3. Select Global Standard as the Deployment type.
  4. Review Tokens per Minute Rate Limit and Guardrails.
  5. Select Deploy. The endpoint is ready when the deployment status is Succeeded.
Replace <fine_tuned_model> with the fine_tuned_model string from the succeeded job. When provisioningState reaches Succeeded, the deployment is ready for inference.

Evaluate the fine-tuned model

You’re responsible for evaluating whether the fine-tuned model meets your application requirements. Evaluation data and methods are separate from the service’s required training inputs. Loss curves alone don’t establish improvement in task performance or clinical quality.
For model-specific approaches and metrics, see Evaluate CxrReportGen Premium and Evaluate MedImageInsight Premium.

Troubleshooting

Clean up resources

When you no longer need a fine-tuned deployment or the files you uploaded for training and validation, delete them: For model-specific hyperparameter information, JSONL schemas, and runnable examples, see: