Publishing9 min readUpdated September 2026

Audio transcription example: three formats, same 40 seconds

By The Hilite team

01 — The example

The same 40 seconds, transcribed three ways

Most people searching for a transcription example want to see one before they commit to a format. Here is a single clip — a supplier call between two people — written out three different ways. The audio is identical in all three. Only the transcription convention changes.

Full verbatim

DANIEL OKONKWO: So the— um, the thing is, we, we quoted you for forty units and you've come back asking for, what, sixty?

MARIA VELEZ: Sixty-two, yeah. [clears throat] Sorry — sixty-two. The extra came in Friday afternoon, like, right at the end of the day.

DANIEL OKONKWO: Right. Right, okay. And is that— I mean, is that a one-off or is that going to be the new, the new normal?

MARIA VELEZ: [sighs] Honestly? I don't know yet. Ask me in a month.

Clean verbatim

DANIEL OKONKWO: The thing is, we quoted you for forty units and you've come back asking for sixty?

MARIA VELEZ: Sixty-two. The extra came in Friday afternoon, right at the end of the day.

DANIEL OKONKWO: And is that a one-off, or is that going to be the new normal?

MARIA VELEZ: Honestly, I don't know yet. Ask me in a month.

Timestamped

[00:02:11] DANIEL OKONKWO: The thing is, we quoted you for forty units and you've come back asking for sixty?

[00:02:18] MARIA VELEZ: Sixty-two. The extra came in Friday afternoon, right at the end of the day.

[00:02:25] DANIEL OKONKWO: And is that a one-off, or is that going to be the new normal?

[00:02:31] MARIA VELEZ: Honestly, I don't know yet. Ask me in a month.

02 — Choosing

Which of the three you actually want

Full verbatim keeps every false start, filler word and audible breath. It exists because sometimes how something was said is the evidence. Use it for legal and disciplinary records, for linguistic or conversation analysis, and for any research where hesitation is data. It is the slowest to read and the most expensive to produce.

Clean verbatim removes fillers, stammers and repeated words but changes nothing else. No sentence is rewritten, no meaning is tidied. This is the default for almost everything: published interviews, podcast transcripts, meeting records, subtitles. Notice that in the example above, Maria's correction from sixty to sixty-two survives the clean pass, because it is a fact, not a stumble.

Timestamped is clean verbatim with a time marker at each speaker turn. Add it when someone will need to find the audio behind a line — a fact-checker, an editor cutting clips, a student revisiting a lecture. Timestamps at every turn are the common convention. Timestamps every thirty seconds regardless of who is speaking is the other, and it is worse, because it puts markers in the middle of sentences.

Tip

If you are not sure, produce clean verbatim with timestamps. You can always strip the timestamps out later. Putting them back in means going through the audio again.

03 — The conventions

Speaker labels, non-speech sounds and the bits people get wrong

Label every turn, never just the first. People arrive at transcripts through search and land in the middle of the page. A label that only makes sense if you read from the top does not work for them.

Use full names in capitals, not initials. DANIEL OKONKWO scans faster than DO, and initials collide the moment a third person joins.

Mark non-speech sounds in square brackets, and only when they carry meaning. Maria's sigh above changes how the line reads, so it stays. Chair noise, a passing siren and the third throat-clear do not, so they go. The test is whether a reader who cannot hear the audio would misread the line without it.

Mark what you could not hear as [inaudible 00:02:44] with the timestamp, rather than guessing. A guess that turns out wrong is worse than a gap, because nobody knows to check it.

Break paragraphs at speaker turns. Inside a long answer, break at natural pauses. Fixed-length paragraphs are the clearest sign of raw machine output that nobody read afterwards.

04 — Machine output

What automatic transcription gets right, and where you still have to look

Automatic transcription is now good enough that the job has changed from typing to checking. On clean single-speaker audio it is close to right. The places it still misses are predictable, which means you can check them deliberately rather than rereading everything.

Proper nouns are the first thing to check: names, companies, product names, places. A model that has never seen a name will produce a plausible-looking wrong one, and plausible is exactly what makes it hard to spot on a reread.

Numbers are the second. Sixty and sixty-two sound alike at speed, and in the example above the difference is the entire point of the conversation.

Overlapping speech is the third. When two people talk at once, most systems attribute the whole stretch to whoever is louder. Check every point where the conversation gets heated.

The video below runs through uploading a file, generating the transcript and downloading it, which is the mechanical half of the job. The checking above is the half that is still yours.

Uploading a file, generating the transcript and exporting it — about a minute end to end.

05 — Before you send it

A short pass before the transcript leaves your hands

Run these in order. It takes a few minutes on a half-hour recording and catches almost everything that would embarrass you later.

  • Every proper noun spelled the way its owner spells it
  • Every number and date checked against the audio
  • Speaker labels on every turn, consistent throughout
  • Overlapping sections listened to again and re-attributed
  • Anything unclear marked [inaudible] with a timestamp, not guessed
  • One format throughout — no half-timestamped documents
  • A note at the top saying which verbatim level this is

Turn any recording into a transcript you can download.

Upload, transcribe, export. Free to try.

Try Hilite free