ModelRefs / Voice Agent — AI Glossary
Voice Agent — AI Glossary
An AI agent combining speech recognition, LLM reasoning, and speech synthesis to conduct spoken conversations in real time. Latency target: <500 ms turn-taking.
Overview
Voice agents (OpenAI Realtime API, ElevenLabs Conversational AI, Vapi, Bland AI) process audio end-to-end with low latency. Key components: STT (Whisper, Deepgram), LLM (GPT-4o audio, Gemini Live), TTS (ElevenLabs, Azure Neural TTS). Latency target: <500 ms turn-taking. Application: customer support automation, voice-first interfaces.
Reference details
| Topic | agents |
|---|---|
| Also known as | conversational voice AI |
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Voice Agent — AI Glossary.
Frequently asked questions
What is Voice Agent?
An AI agent combining speech recognition, LLM reasoning, and speech synthesis to conduct spoken conversations in real time.
Is Voice Agent the same as conversational voice AI?
Yes — conversational voice AI are common aliases for Voice Agent.
What concepts are related to Voice Agent?
Closely related concepts include agent, tool agent, conversational ai.