Transcription used to be the bottleneck. It is not any more — accuracy on clear English audio is high across every serious tool, much of it runs on-device, and the cost is close to zero.
Which means the interesting question has changed. It is no longer how do I get a transcript. It is what do I do with an hour of text that nobody, including me, is going to read.
The three outputs
Any recording can become one of three things, and choosing deliberately is most of the skill:
| Output | When | Effort |
|---|---|---|
| A full transcript | You need exact words — quotes, interviews, legal, research | Automatic |
| A summary with timestamps | You need the gist and the ability to jump back | Automatic, then a read-through |
| Three lines you wrote | Almost everything else | Two minutes of yours |
The default should be the third, and almost everyone defaults to the first because it is free. A transcript feels like a record and functions as an unread archive — the same pattern as saving articles you never open, which is the collector's fallacy with better technology.
How to actually get one
On your phone, already built in. Apple's Voice Memos and Notes offer transcription on-device on recent hardware, and Android has equivalents in Recorder. For a personal memo this is enough and nothing leaves the device — which matters for anything sensitive.
Whisper, locally. OpenAI's open-source model runs on a normal laptop through several front-ends. It is the best option when the audio must not go to a service: client material, medical, legal, personal. Slower than cloud, free, and private.
Cloud services (Otter, Descript and the rest) add speaker labels, editing and search across a library. Speaker separation is the real reason to use them for multi-person recordings — it is the thing local tools do worst.
Descript is worth knowing about specifically because it inverts the workflow: edit the transcript and the audio edits with it, which is genuinely useful if the audio itself is the deliverable.
Where accuracy still fails
The "95% accurate" figure is for clear, single-speaker, native-accented English. It degrades, sometimes sharply, on:
- Crosstalk. Two people talking over each other, which is most real conversation.
- Accents and non-native speakers, unevenly and sometimes unfairly.
- Technical vocabulary and proper nouns. Names, drugs, products, jargon — precisely the words you needed.
- Poor audio. A phone in a pocket, a café, a car.
The dangerous part is that errors read as confident text. A transcript that says the wrong number, or renames a person, does not look wrong. Never quote from an unverified transcript — listen back to any passage you intend to rely on.
The processing step nobody does
A transcript is not a note. It is a searchable recording, and it has the same problem the recording had: you have to go through it to find anything.
The step that makes it useful takes two minutes:
- Skim the transcript once, the same day.
- Write three to five lines in your own words: what this was, what mattered, what to do.
- Keep the transcript underneath if exact wording might matter, and delete it if not.
Writing those lines yourself, rather than generating them, is the difference between remembering the content and owning a file about it — see the generation effect. An AI summary is fine as a starting point and it is not a substitute, because the summariser does not know which sentence mattered ten times more than the rest.
What to leave as audio
Not everything should become text.
- Musical ideas. Text cannot hold a melody — see best notes app for musicians.
- Tone that matters. How someone said something, for a difficult conversation.
- Anything where you will want the original. Keep the audio and the transcript together; storage is cheap.
Conversely, transcribe anything you will want to search. A year of voice memos is unsearchable as audio and fully searchable as text, and that alone justifies transcription even if you never read a word of it.
The privacy question
Sending audio to a cloud service means sending everything on it — including the other people in the room, who did not choose that. Before uploading, ask what the service retains, for how long, and whether it trains on your data.
For client work, medical, legal or anything under NDA, use on-device or local transcription. This is one of the clearest cases where local processing is not paranoia but a requirement. See encrypted notes.
And recording other people carries consent obligations that vary by jurisdiction — see how to take notes on a phone call.
The minimum viable version
Record. Let your phone transcribe it. Skim it once that day and write three lines of your own. Delete the ones that were nothing.
The three lines are the note. Everything else is the raw material — and raw material nobody processes is just storage.
More in voice note-taking, how to take notes from a video, and how to take notes.