01 — The options
Three ways to do it, and what each actually costs
Automatic transcription runs the audio through a speech model and hands you text in roughly the time it takes to upload the file. On clear single-speaker audio it is close to right. On a noisy four-person meeting it is not. Cost is measured in pennies or included in a subscription.
Manual transcription means a person listens and types. Reckon on four to six hours of work per hour of clear audio if you have never done it, and considerably more for overlapping speech or an unfamiliar accent. Almost nobody should choose this in 2026, because the hybrid below gets you the same accuracy for a fraction of the effort.
Hybrid is the answer for nearly every real job: generate automatically, then check the four places errors cluster. On a one-hour recording the checking pass takes twenty to forty minutes, not four hours, because you are reading against the audio rather than typing from it.
The exception is audio where an error has consequences — court records, regulated clinical work, anything that will be challenged. There, a human transcription service that guarantees accuracy is worth the per-minute cost, because the guarantee is what you are buying.
Decide which of the three you are doing before you start, not after you see the machine output. People who skip the checking pass because the draft looked fine are the ones who find the error in public.
02 — Upstream
The transcript is mostly decided before you press record
Every minute spent on capture saves several in checking. A speech model cannot separate two voices that arrived on one channel at the same volume, and it cannot recover a word that was lost under an air conditioner.
Get the microphone close. A phone on the table between two people produces far better source audio than a laptop microphone across a room, and better audio produces a better transcript before any software is involved.
Record each speaker to a separate track where you can. Separate tracks make speaker attribution almost trivial; a single mixed track makes it guesswork the moment two people overlap.
Ask everyone to say their name at the top. Fifteen seconds, and it gives you a clean reference sample for every ambiguous attribution later.
Kill the obvious noise sources before you start rather than trying to remove them afterwards. Fan, air conditioning, a window onto a road. Noise removal always costs some of the voice along with the noise.
03 — Doing it
Upload, generate, download
The mechanical part is genuinely short now: put the file in, wait, take the text out. The video below runs the whole thing end to end in under a minute, including downloading the finished transcript.
Two things worth deciding before you generate. First, whether you want timestamps — adding them later means going back through the audio, so turn them on now even if you strip them out afterwards. Second, what format you need out: plain text for reading, SRT or VTT if the transcript is going to become captions.
If the audio is long, transcribe it in one piece rather than splitting it. Splitting introduces boundary errors exactly where you cut, and makes speaker labels restart from scratch in each segment.
04 — The checking pass
The four places errors actually live
Do not reread the whole transcript. Automatic errors are not spread evenly, they cluster, and checking the clusters takes a fraction of the time for nearly all of the benefit.
Proper nouns. Names of people, companies, places and products. A model that has not seen a name produces a confident, plausible, wrong one — and plausible is what makes it survive a reread. Search for each name once and fix every instance together.
Numbers. Prices, dates, quantities, percentages. Fifteen and fifty sound alike at conversational speed, and a number is usually the part of a sentence someone will act on.
Overlaps. Find every turn shorter than about three seconds, or any that ends mid-sentence, and listen to those spots. That is where two people talked at once and the system gave the whole stretch to whoever was louder.
Domain vocabulary. Jargon, acronyms and technical terms specific to the field. If the recording is about mitral valves or mezzanine financing, assume the model has guessed at least once.
Anything you genuinely cannot make out gets marked [inaudible] with its timestamp. A marked gap is honest and checkable. A confident guess is neither.
05 — Afterwards
What to do with the transcript once you have it
For a published episode or video, the transcript is also the captions, the show notes source and the thing search engines can actually read. Exporting SRT at the same time costs nothing extra and saves doing it again later.
For research or journalism, store the transcript and the audio together with a header recording the date, the participants, the verbatim level and who checked it. In six months none of that will be in anyone's memory.
If the recording is sensitive, check what the transcription service retains and for how long before you upload, rather than afterwards. Retention settings are a policy decision, not a technical detail.
Upload a recording and get the text back.
Transcribe, download, and edit the audio in the same place. Free to try.