Automate Basics

Glossary

What is AI transcription?

Transcription is the conversion of spoken audio, such as a meeting, call, or interview, into written text, which AI tools can do automatically.

AI transcription uses speech recognition models to turn recorded or live speech into text. Many tools also label who said what, sometimes called speaker diarization, add timestamps so readers can jump to a moment in the recording, and make the text searchable later. Meeting assistants build on transcription to produce summaries, action items, and answers to questions about what was said.

For example, a team might use a meeting assistant that joins video calls, records them with everyone's knowledge, and produces a transcript with a short summary and a list of decisions and follow-ups. Someone who missed the meeting can read the summary, then search the transcript for the moment a particular client or deadline was mentioned.

Accuracy depends heavily on the audio. Clear speech from one person on a good microphone usually transcribes well, while people talking over each other, background noise, accents the model has had little exposure to, and specialist terms, names, and acronyms cause errors. Numbers and names are especially easy to get wrong, so they should be checked in any transcript used for records, quotes, or decisions.

Recording and transcribing conversations raises consent and privacy questions. Laws and company policies often require telling participants, or getting their agreement, before a call is recorded, and transcripts may contain confidential or personal information. It is worth checking where transcripts are stored, who can see them, and how long they are kept.

An example

After a client call, the meeting assistant's transcript lets a project manager copy the exact wording of the agreed deadline into the project plan.

Tools where you will meet it