Glossary

What is speech-to-text?

Definition

Speech-to-text (STT), also called speech recognition, converts spoken audio into written text. It's the first step in a voice AI system: turning what a caller says into text a language model can understand.

01How it's used in phone AI

On a call, speech-to-text transcribes the caller's words in real time so the system can interpret intent. Accuracy with accents, background noise, and industry terms directly affects how well the AI understands and responds.

Frequently asked questions

How accurate is speech-to-text on phone calls?

Phone audio is harder than a quiet dictation setting — it's compressed, often noisy, and callers interrupt themselves. Modern engines handle typical calls well, but accuracy drops with heavy background noise, crosstalk, and unfamiliar names or jargon. Systems built for telephony, and ones that can learn your industry's vocabulary, close much of that gap. There is no single accuracy number; it depends on the audio.

What's the difference between speech-to-text and voice recognition?

Speech-to-text figures out what was said, producing a transcript. Voice recognition — more precisely, speaker recognition — figures out who is speaking, often for authentication. The terms get mixed up because both listen to speech, but they answer different questions. A phone AI system relies on speech-to-text; speaker identification is a separate, optional layer.

Does speech-to-text work in real time during a call?

Yes. Voice AI systems use streaming speech-to-text, which transcribes audio as the caller speaks instead of waiting for them to finish. That's what lets the assistant start formulating a response immediately and reply without an awkward pause. Batch transcription — processing a whole recording afterward — is the slower cousin used for call summaries and records.

Related terms

Ahoya is an AI receptionist that answers every call 24/7.

Start free