Skip to content
Sarudo

Voice Providers & Options

How Sarudo's voice features work — which options are live, which are optional add-ons, and which are on the roadmap.

providerssttttsinbound calls

The Voice Stack at a Glance

Sarudo's voice feature composes three independent layers: telephony (placing and receiving the call), AI voice (real-time conversation with the person on the other end), and speech-to-text (transcribing the recording after the call). Each layer has one primary provider wired today plus optional alternatives you can opt into if you need more voice variety, better accuracy on specialized vocabulary, or lower latency. The defaults are chosen to get a fresh client talking to their AI employee over the phone with minimal setup.

The defaults (Twilio telephony + AI voice + on-device transcription) are enough for most business outbound calling. The alternatives below are opt-in — no reason to touch them unless you have a concrete need.

Telephony — Twilio

Telephony — placing the actual phone call and handling call routing — is Twilio. You provision a Twilio account and a phone number during setup, and your AI employee places calls using your number as caller ID. Twilio is the only telephony provider wired today; there is no planned alternative, because Twilio is effectively the industry standard and swapping it would not buy anything meaningful. Per-minute charges bill to your Twilio account at Twilio's rates, with nothing added on top.

AI Voice

The AI voice layer handles the real-time conversation — the "your AI employee talking to the other person" part. It generates the natural-sounding voice responses on the fly as the call happens, using your Twilio number as the caller ID. An alternative AI voice engine is also available and can be swapped in by your setup team if you specifically need a lower-latency cold-call feel or a particular set of voice personas. If AI voice is not configured, calls fall back to a basic built-in text-to-speech — functional, but clearly robotic and not recommended for real client-facing calls.

An alternative AI voice engine is available but not yet routed by default. If you want to switch your AI voice to it, let your setup team know — it is a configuration change, not a development effort.

Speech-to-Text — On-Device (primary)

Post-call transcription runs on an on-device speech-to-text engine on your dedicated server, which means audio never leaves your infrastructure. This is the only transcription option wired today, and it is typically the right choice — quality is strong for clear business speech, latency is the processing time alone (no network round-trip), and there is no per-minute transcription bill. If you have a particular need for a cloud-based transcription tier (extremely long calls where server-side processing is slow, or specialized medical/legal vocabularies), higher-accuracy cloud options can be turned on by your setup team.

TTS Alternatives (when AI voice is not live)

When the AI voice is not handling the conversation — for a simple reminder where you want a voice playback of fixed text rather than an interactive call — Sarudo can generate that playback in a range of voices. A built-in voice runs on your own server and costs nothing extra; higher-quality voices from a paid provider can be turned on during setup, and those run on an account in your name and bill per character used. The current voice-call pipeline is built around the interactive path, so fixed-text playback only comes into play when you ask for it — say "send a voice reminder to John saying [text]" and your AI employee picks the best available voice for the job.

When generating a fixed-text voice message, Sarudo automatically uses the best available voice quality turned on for your instance. You can request a specific voice profile if you want a particular sound for a given message.

Inbound Calls

Outbound calling is fully live — your AI employee can place calls on your behalf and is designed around that primary use case. Inbound calling is partially wired: Twilio can forward inbound calls on your provisioned number to the AI voice layer, which can handle a real-time conversation. What is not yet live is the full inbound flow — call routing logic (who the AI voice should be when answering, what business context should it know, how should it escalate to you), voicemail handling, and after-call logging into your CRM. If inbound answering is a hard requirement for your use case, it is feasible to enable manually during onboarding; if you are treating voice primarily as an outbound tool (most common), the defaults are already what you want.

If you publish your Twilio number anywhere and expect calls to come in, talk to your setup team before going live. The inbound path is configurable but is not the default, and you will want to decide the behavior explicitly.