ModelRefs / Voice AI — Canonical Workflow

Voice AI — Canonical Workflow

Canonical Voice AI workflow: ASR, LLM, TTS, latency budgets and deployment patterns.

Overview

Voice AI combines speech-to-text, an LLM and text-to-speech behind a low-latency turn-taking pipeline. The canonical stack budgets every hop for sub-second response so conversations feel natural over the phone or in real-time apps.

Implementation profile

Categoryaudio-models
Implementation maturityproduction
Evidence statusincomplete
Primary use casescustomer-support
Deployment optionsmanaged-api, edge
Architecturesserverless-api, edge-runtime

Candidate models with published references

Coverage means the model is a candidate worth evaluating for this workflow, not a ranking or a recommendation. Models whose reference pages are still in review are omitted.

Benchmarks relevant to this workflow

swe-bench, gpqa, mmlu, simpleqa, mgsm, mmlu-pro.

Relevance is a coverage signal from the canonical registry. Each benchmark only describes its own protocol and date, so confirm the harness matches your workload before treating a score as evidence.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Voice AI — Canonical Workflow.