ModelRefs / Voice Agent — AI Glossary

Voice Agent — AI Glossary

An AI agent combining speech recognition, LLM reasoning, and speech synthesis to conduct spoken conversations in real time. Latency target: <500 ms turn-taking.

Overview

Voice agents (OpenAI Realtime API, ElevenLabs Conversational AI, Vapi, Bland AI) process audio end-to-end with low latency. Key components: STT (Whisper, Deepgram), LLM (GPT-4o audio, Gemini Live), TTS (ElevenLabs, Azure Neural TTS). Latency target: <500 ms turn-taking. Application: customer support automation, voice-first interfaces.

Reference details

Topicagents
Also known asconversational voice AI
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Voice Agent — AI Glossary.

Frequently asked questions

What is Voice Agent?

An AI agent combining speech recognition, LLM reasoning, and speech synthesis to conduct spoken conversations in real time.

Is Voice Agent the same as conversational voice AI?

Yes — conversational voice AI are common aliases for Voice Agent.

What concepts are related to Voice Agent?

Closely related concepts include agent, tool agent, conversational ai.