Skip to main content
In this article, you compare Voice Live with Foundry Agent Service with voice-based prompt agents in Microsoft Foundry. Then you migrate your voice experience while keeping the existing text agent available. Here, Voice Live integration means Voice Live connected to an existing Foundry text agent. A Foundry voice agent is a separate agent with kind: voice. It still uses Voice Live for the voice runtime, but you manage its model, instructions, audio, and tools as an agent definition. Migration creates a new agent alongside the original. You can’t change an existing text agent’s interaction mode to voice. This process is separate from migrating Agent Service (classic) to the new Agent Service.
Items marked preview in this article are currently in preview. This preview is provided without a service-level agreement, and Microsoft doesn’t recommend it for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.

Prerequisites

  • A working Voice Live integration with a Foundry text agent, including access to its configuration and client code.
  • A Foundry project with access to the voice-agent preview and a supported model in the project’s region.
  • Foundry User role on the target project, with permission to create and invoke agents.
The Foundry RBAC roles were recently renamed. Foundry User, Foundry Owner, Foundry Account Owner, and Foundry Project Manager were previously named Azure AI User, Azure AI Owner, Azure AI Account Owner, and Azure AI Project Manager. You might still see the previous names in some places while the rename rolls out. The role IDs and core permissions are unchanged by the rename.
  • Access to the tool connections and data resources that the migrated agent needs. Have a resource owner grant any required permissions.
  • A microphone and speakers or a headset for testing spoken conversations.
  • For the Python connection example, Python 3.10 or later and the Azure CLI, signed in with az login.

Compare the two approaches

Compare the models you can use, who manages them, and how you operate the voice experience. Both approaches provide configurable voices, turn detection, and interruption handling through Voice Live. The Voice Live column refers only to the Foundry text-agent integration. Voice Live used directly also supports speech-to-speech models. Choose a Foundry voice agent for speech-to-speech models, service-managed model hosting, built-in telephony, or integrated voice configuration and call review. You can also keep a supported text-model approach; migration doesn’t require switching to gpt-realtime. Model, region, network, and transport restrictions still apply. Measure conversation quality, latency, and cost with your workload before switching.

Choose how to reuse your existing agent

Choose the conversation design before copying configuration. Reusing a text agent as a specialist is different from forwarding every spoken turn to it. For the subagent approach, the subagent must be in the same project as the voice agent. Subagents aren’t supported in VNet-isolated projects. Follow Use a subagent in a voice-based agent and test delegation, follow-up questions, and failure behavior. The remaining steps focus on a model-backed voice agent, with optional subagents. A custom hosted voice application that exposes invocations_ws is a different integration; see Build a voice agent with hosted agents.

Inventory the source and create a new agent

Capture the effective configuration, not just the values stored on the text agent.
  1. Record the source agent name and version, model deployment, instructions, tools, and project connections.
  2. Export the Voice Live settings. If you used the integration sample, reassemble microsoft.voice-live.configuration and its numbered metadata entries, such as .1 and .2, in order.
  3. Record any client-side session.update overrides, proactive greeting code, audio formats, interruption handling, and channel configuration.
  4. Record authentication and cross-resource dependencies, including any foundry_resource_override and managed identity settings. Don’t copy credentials into the new agent definition.
  5. Follow Create a voice-based prompt agent to create an agent with a new name. In the portal, select Voice as the interaction mode. In code, use VoiceAgentDefinition instead of PromptAgentDefinition.
  6. Choose a supported speech-to-speech model, such as gpt-realtime, or a supported text model. Both model families support service-managed and self-deployed options.
  7. Set model_type to managed for a service-managed model, or self_deployed for your own supported deployment. Don’t treat a deployment name as a managed model name.
  8. Adapt the instructions for spoken interaction: short answers, one question at a time, and explicit confirmation before consequential actions. Preserve the source agent’s business rules and escalation requirements.
  9. Save the new agent version and record its name and version for testing. Keep the original text agent and client configuration unchanged until the new path is ready.
For SDK creation examples, use the explicit-definition quickstart. Creating a voice agent version doesn’t copy the source agent’s tools, connections, or conversation history.

Move voice settings into the definition

Translate the settings rather than copying the old session object into metadata. The following table uses the JSON field names from the Voice Live integration samples and the voice-agent definition. For the complete schema and SDK examples, see Configure a voice agent. SDK and azure.yaml property names can differ from the JSON names in this table. The service applies the definition when a session starts. Remove client startup code that resends the full legacy session configuration. Keep only supported per-session overrides, such as audio or avatar settings, and verify the effective configuration returned by the service.

Reconnect tools and knowledge

Voice agents provide the same tool and knowledge coverage as text agents, but configuration and execution paths differ. Map the existing integrations to the voice-agent tool interfaces instead of copying the text agent’s tools array unchanged. Confirm which identity accesses each dependency.
  • For a native function tool, configure its schema on the voice agent and implement the function-call response in your live client. A schema alone doesn’t execute the function.
  • For an MCP tool, configure the server and the required project connection or authentication. Verify access from the target project.
  • For Foundry server-side tools, such as Azure AI Search or OpenAPI tools, create or reuse a supported toolbox and reference its name and version.
  • For a retained text subagent, configure its name, capabilities, and optional version. Test which requests the voice agent answers directly and which it delegates.
  • Test successful calls, denied permissions, timeouts, and unavailable dependencies before moving user traffic.
See Attach tools for configuration details and the client-executed function sample for the live tool-call loop.

Update the client connection

Use the Foundry project endpoint, not the Voice Live resource endpoint. The project endpoint has this form: https://<account>.services.ai.azure.com/api/projects/<project>. The voice-agent API uses api-version=v1, but voice agents remain a preview feature. REST requests require the Foundry-Features: VoiceAgents=V1Preview opt-in. The Python realtime helper supplies this header automatically. For a raw WebSocket client, use wss:// and replace /voice-live/realtime with /api/projects/<project>/agents/<agent-name>/endpoint/protocols/voice?api-version=v1. Request a Microsoft Entra token for https://ai.azure.com/.default, send it in the Authorization header, and include the voice-agent preview opt-in. Resource API keys aren’t a substitute for agent authentication. Don’t put bearer tokens in the URL or long-lived credentials in browser code. The project and agent are now part of the endpoint path. Review any old cross-resource connection parameters separately rather than appending them to the new URL.

Check the connection with Python

In Python, replace the azure-ai-voicelive live-session connection with azure-ai-projects and beta.voice_agents.realtime. Install the Azure AI Projects client library with its voice dependencies:
Reference: Azure AI Projects client library for Python. Set these environment variables in the shell that runs the example: Save the following example as check_voice_connection.py, and run it with python check_voice_connection.py. It opens a session against that exact version, waits for configuration to complete, and then closes the connection. It doesn’t send microphone audio.
Reference: Python realtime client and event models. The expected result is Voice session ready.. For a complete spoken turn, follow Talk to the agent or adapt the live microphone sample. Use the current Projects SDK samples for the public preview. Earlier private-preview samples that use a separate azure-ai-voiceagents package or /voice_agents management route don’t describe this API.

Adapt audio and event handling

The new endpoint isn’t just a URL replacement for an older Voice Live API. For raw JSON clients, update these event families, including their delta and done events: The Projects SDK maps these events to typed models. For example, RealtimeServerEventResponseAudioDelta represents response.output_audio.delta, and its delta contains decoded audio bytes. Before enabling microphone input, wait for session.updated. Match the input and output formats to the saved definition. Preserve local playback-buffer clearing when the caller interrupts; canceling generation alone doesn’t clear audio already queued on the device. If the definition includes a greeting, handle its response separately from the reply to the first user turn. Don’t send the old proactive greeting as well. WebRTC isn’t a drop-in substitute for WebSocket. The current voice-agent WebRTC transport doesn’t support self-deployed models or hosted conversation engines. Use WebSocket for those configurations and validate each required browser or telephony path separately.

Configure history, recordings, and monitoring

Treat the new voice conversation as a new record, not as a continuation of the old text-agent conversation.
  • Start a new voice conversation at cutover. Retain prior text-agent history according to your existing policy; don’t pass its conversation ID expecting an automatic import.
  • Decide whether to enable store. Its default is false. Setting it to true stores the transcript, event timeline, and raw audio together, so review notice, consent, access, and retention requirements first.
  • Capture the voice conversation ID from session.created for correlation and, when storage is enabled, read-back through beta.voice_agents.conversations. Don’t confuse it with a realtime session ID.
  • Retrieve merged recordings after the session ends. Recording finalization isn’t necessarily immediate; follow the conversation audio sample.
  • Configure voice-agent tracing and monitoring. Conversation storage and sensitive-content capture in traces are separate controls.
Review voice-agent pricing for the model-hosting choice, audio usage, tools, and optional features. Measure the migrated workload rather than assuming unchanged latency or cost.

Validate and switch traffic

Test the new agent alongside the existing integration before changing production routing. Follow these steps in order to switch traffic:
  1. Pin the tested voice-agent version and any toolbox or subagent versions your release depends on.
  2. Update a test client or a limited traffic route to use the new voice endpoint.
  3. Configure the required voice-agent channels. Existing text-agent app publishing or phone configuration doesn’t migrate merely because you create a voice agent.
  4. Increase traffic only after the checks pass. Let existing calls finish on their original integration rather than changing the endpoint during a call.
  5. Keep the old endpoint, agent, and client configuration available for rollback. If needed, route new sessions back to the original integration.
  6. Retire the old path only after confirming that no applications still depend on it and that required historical records are retained.

Troubleshoot migration issues

Use the failing stage to narrow the configuration difference.