Voice Providers & Options
How Sarudo's voice features work — which options are live, which are optional add-ons, and which are on the roadmap.
The Voice Stack at a Glance
Sarudo's voice feature composes three independent layers: telephony (placing and receiving the call), AI voice (real-time conversation with the person on the other end), and speech-to-text (transcribing the recording after the call). Each layer has one primary provider wired today plus optional alternatives you can opt into if you need more voice variety, better accuracy on specialized vocabulary, or lower latency. The defaults are chosen to get a fresh client talking to their AI employee over the phone with minimal setup.
The defaults (Twilio telephony + AI voice + on-device transcription) are enough for most business outbound calling. The alternatives below are opt-in — no reason to touch them unless you have a concrete need.
Telephony — Twilio
Telephony — placing the actual phone call and handling call routing — is Twilio. You provision a Twilio account and a phone number during onboarding, and your AI employee places calls using your number as caller ID. Twilio is the only telephony provider wired today; there is no planned alternative because Twilio is effectively the industry standard and swapping it would not buy anything meaningful. Per-minute charges bill to your Twilio account, not through your Sarudo subscription.
AI Voice
The AI voice layer handles the real-time conversation — the "your AI employee talking to the other person" part. It generates the natural-sounding voice responses on the fly as the call happens, using your Twilio number as the caller ID. An alternative AI voice engine is also available and can be swapped in by your setup team if you specifically need a lower-latency cold-call feel or a particular set of voice personas. If AI voice is not configured, calls fall back to a basic built-in text-to-speech — functional, but clearly robotic and not recommended for real client-facing calls.
An alternative AI voice engine is available but not yet routed by default. If you want to switch your AI voice to it, let your setup team know — it is a configuration change, not a development effort.
Speech-to-Text — On-Device (primary)
Post-call transcription runs on an on-device speech-to-text engine on your dedicated server, which means audio never leaves your infrastructure. This is the only transcription option wired today, and it is typically the right choice — quality is strong for clear business speech, latency is the processing time alone (no network round-trip), and there is no per-minute transcription bill. If you have a particular need for a cloud-based transcription tier (extremely long calls where server-side processing is slow, or specialized medical/legal vocabularies), higher-accuracy cloud options can be turned on by your setup team.
TTS Alternatives (when AI voice is not live)
When the AI voice is not handling the conversation (for example, for simple reminder messages where you just want a voice playback of a fixed text rather than an interactive conversation), Sarudo can generate the voice playback in a range of voices. A free built-in voice option is bundled, and higher-quality premium voices can be enabled by your setup team. The current Sarudo voice-call pipeline focuses on the interactive AI-voice path, so these fixed-text playback options come into play only when you explicitly ask for one — ask your AI employee "send a voice reminder to John saying [text]" and it picks the best available voice for the job.
When generating a fixed-text voice message, Sarudo automatically uses the best available voice quality your plan has enabled. You can request a specific voice profile if you want a particular sound for a given message.
Inbound Calls
Outbound calling is fully live — your AI employee can place calls on your behalf and is designed around that primary use case. Inbound calling is partially wired: Twilio can forward inbound calls on your provisioned number to the AI voice layer, which can handle a real-time conversation. What is not yet live is the full inbound flow — call routing logic (who the AI voice should be when answering, what business context should it know, how should it escalate to you), voicemail handling, and after-call logging into your CRM. If inbound answering is a hard requirement for your use case, it is feasible to enable manually during onboarding; if you are treating voice primarily as an outbound tool (most common), the defaults are already what you want.
If you publish your Twilio number anywhere and expect calls to come in, talk to your setup team before going live. The inbound path is configurable but is not the default, and you will want to decide the behavior explicitly.