Voice input on your website widget, a neural voice answering your phone line — grounded in the same knowledge base as every text channel, with text transcripts and zero stored audio.
Connect a Twilio number and the same agent answers calls out loud: Twilio's speech recognition hears the caller, the agent retrieves from your knowledge base, and a neural Amazon Polly voice speaks the reply — deliberately short sentences, no URLs or list formatting, because the agent knows it's on a call. Inbound only, and every call becomes a text transcript in your dashboard. Phone integration in detail →
No — you enable it per agent: text only, voice only, or text + voice. Until you flip it, the widget is a standard text chat.
On the website widget: OpenAI Whisper for speech-to-text and OpenAI’s TTS for the spoken reply, with six voices to pick from. On the phone: Twilio’s built-in speech recognition and a neural Amazon Polly voice. Two surfaces, two engine sets, both configured for you.
A spoken turn counts as one message against your quota — the transcription and speech steps aren’t double-billed as messages.
No. Widget audio is transcribed and discarded; the spoken reply is generated on the fly and not saved; phone calls are never recorded. Only text transcripts are kept, in your dashboard.
Browsers only expose the microphone in a secure context. Your site is on HTTPS anyway; the widget checks and explains rather than failing silently.
10 minutes from sign-up to your first answered customer message. Free forever plan, no card required to build.