Voice AI

Customers talk. Your agent talks back.

Voice input on your website widget, a neural voice answering your phone line — grounded in the same knowledge base as every text channel, with text transcripts and zero stored audio.

Voice on your website

  • Hold to talk: visitors press the mic, speak, release — the widget transcribes (OpenAI Whisper) and the agent answers as usual.
  • Spoken replies: optionally the agent reads its answer aloud (OpenAI TTS, six voices: alloy, echo, fable, onyx, nova, shimmer), with auto-speak configurable per agent.
  • Barge-in: the moment a visitor starts talking, playback stops — no waiting out a paragraph.
  • Sane limits: recordings cap at ~60 seconds, spoken text is cleaned of citations and links before synthesis, and voice traffic is rate-limited like any other widget traffic.
  • Per-agent switch: text, voice, or text + voice — default is text, so nothing changes until you turn it on.

Voice on the phone

Connect a Twilio number and the same agent answers calls out loud: Twilio's speech recognition hears the caller, the agent retrieves from your knowledge base, and a neural Amazon Polly voice speaks the reply — deliberately short sentences, no URLs or list formatting, because the agent knows it's on a call. Inbound only, and every call becomes a text transcript in your dashboard. Phone integration in detail →

Privacy: transcripts yes, recordings never

  • No audio recording on any surface — widget clips are transcribed and discarded, phone calls are never recorded.
  • Generated speech is streamed to the listener and not stored.
  • What is kept: the text transcript of each conversation, visible to your team in the dashboard — consistent with our phone no-recording policy.

Frequently Asked Questions

Is voice on by default?

No — you enable it per agent: text only, voice only, or text + voice. Until you flip it, the widget is a standard text chat.

Which engines power it?

On the website widget: OpenAI Whisper for speech-to-text and OpenAI’s TTS for the spoken reply, with six voices to pick from. On the phone: Twilio’s built-in speech recognition and a neural Amazon Polly voice. Two surfaces, two engine sets, both configured for you.

Do voice messages cost extra messages?

A spoken turn counts as one message against your quota — the transcription and speech steps aren’t double-billed as messages.

Is any audio stored?

No. Widget audio is transcribed and discarded; the spoken reply is generated on the fly and not saved; phone calls are never recorded. Only text transcripts are kept, in your dashboard.

Why does the mic need HTTPS?

Browsers only expose the microphone in a secure context. Your site is on HTTPS anyway; the widget checks and explains rather than failing silently.

Give your support a voice. Literally.

10 minutes from sign-up to your first answered customer message. Free forever plan, no card required to build.