Can ChatGPT transcribe an interview?
It will give you the words, and it will not reliably tell you who said them. For an interview that is the one thing you needed. The base model has no speaker diarization at all, so a two-person conversation comes back as an undifferentiated block of speech, and attributing a quote means listening to the whole recording anyway. That, plus the file size cap on anything over an hour, is why most interviewers give up on this route.

Upload the recording, click any line to hear who said it.
Transcribing an interview with the speakers kept straight
- 1
Record the speakers separately if you can
Two separate tracks is the one thing that makes speaker attribution trivial in any tool. If the interview has not happened yet, this is the highest-value decision available to you.
- 2
Split the file
A one-hour interview at ordinary quality will be too large for a single request. Splitting it means the transcript arrives in pieces, and whatever speaker guesses the model made in part one will not carry into part two.
- 3
Label the speakers yourself
Expect to go through the returned text marking who is who. On a long interview this is the bulk of the work, and it is work you are doing because the tool did not.
- 4
Verify every quote you plan to publish
This is the part that matters. A misattributed quote is a correction, not a typo. Without a transcript linked to the audio there is no fast way to check, so you end up listening again.

Keep working inside the Hilite editor
Your uploaded files open directly in a focused editor view. Arrange your clips, preview the timeline, adjust transitions, and export when the track sounds perfectly right.
- Review the transcript in context
- Keep timestamps with the words
- Copy, export, or keep editing
Input formats
Output formats
Built for every kind of audio
Whatever you are making, this tool helps you get to a cleaner, more useful result, then hand off to Hilite to enhance and polish.
- An interview is speakers, not just speechThe transcript of a monologue is text. The transcript of an interview is text plus attribution. Drop the attribution and you have not transcribed the interview, you have transcribed the sound.
- Verification is the real costJournalists, researchers and podcasters all spend more time checking quotes than reading transcripts. A transcript where clicking a line plays that moment collapses that cost to nearly nothing.
- Crosstalk is where it breaksPeople interrupt each other. That is what makes an interview good and what makes it hard to transcribe. Any tool that guesses at speakers guesses worst exactly where the conversation is liveliest.

Made for podcasters and their teams
Trim the dead air off the top and tail, and every recording is ready to edit, enhance, and publish.






FAQ
Common questions
If you do not see your question here, the support team replies within one business day.
- Not dependably. The base Whisper model has no speaker diarization, and where labels do appear they are frequently swapped or too generic to rely on. Plan to assign speakers yourself.
- Less than an hour in practice. The single-request limit is 25MB and a one-hour MP3 at 128 kbps is already larger than that. Longer interviews have to be cut into parts.
- The word accuracy on clear speech is good. The risk is not spelling, it is attribution and the handful of words the model guesses at. Anything you publish should be checked against the recording, and ChatGPT gives you no quick way to do that.
- Accents are handled reasonably well. Crosstalk is where every transcription engine degrades, and it degrades worst precisely when two people are talking over each other, which in an interview is often the interesting part.
- Speaker labels, timestamps you can click to hear the moment, 99+ languages, and export as TXT, DOCX or SRT. Long interviews go in whole with no splitting. Free to start in the browser.
Transcribe the interview, speakers and all
Upload the whole recording, get speaker labels and timestamps, and check any quote with one click.