How to transcribe a job interview: methods for phone, video and in-person
On this page
- Live transcription versus transcribing a recording afterward
- Video interviews: the recording already exists somewhere
- Phone screens: no platform to lean on
- In-person interviews: the one nobody's platform records
- What determines accuracy, regardless of method
- When automatic transcription is not accurate enough
- File formats, and what to do with them afterward
- Where a live question guide changes the workflow
- Questions people ask
Transcribing a job interview means turning the spoken conversation into text you can search, quote and attach to a scorecard, and there are three ways to get there: live transcription while the call runs, transcription of a recording after the fact, or a human transcription service for interviews where automatic accuracy is not good enough. Which one fits depends on the interview format — phone, video or in person — more than on which tool you already have installed.
This guide covers the methods themselves, independent of any one platform. For the settings on a specific video platform, see our guides to Google Meet, Zoom and Microsoft Teams. This one is for the question those pages don't answer on their own: which method to use, on which format, and what to do with the transcript once you have it.
Live transcription versus transcribing a recording afterward
Live transcription runs while the interview happens, converting speech to text in real time on your screen. Transcribing afterward means recording the audio first and running it through a transcription tool once the call is over. Both end with a written transcript; they differ in what you get during the call.
| Live transcription | Transcribe afterward | |
|---|---|---|
| What you see during the call | Scrolling text, and a question guide you can tick off as it's covered | Nothing extra; you're just talking and listening |
| Accuracy | Slightly lower on cross-talk and fast exchanges | Slightly higher; processing isn't racing the conversation |
| Best for | Structured interviews where you're tracking coverage of must-ask questions | Interviews where you want to stay fully present and review later |
| Failure mode | A dropped connection loses the live feed, not necessarily the recording | If the recording itself fails, there's nothing to transcribe |
Most recruiters land on live transcription for structured interviews with a question bank to cover, and after-the-fact transcription for exploratory conversations, like an intake call, where there is no checklist to track.
Video interviews: the recording already exists somewhere
For Zoom, Teams and Google Meet, transcription usually piggybacks on a recording setting that is already part of the platform, plus a separate transcription toggle that has to be turned on alongside it. The three platforms differ enough in defaults, storage location and what the candidate is shown that they are worth checking individually rather than assumed to behave the same way — see the platform-specific guides linked above for the current settings on each.
What is common across all three: transcription needs a decent audio path more than it needs a good camera. A candidate on a laptop microphone from across a noisy kitchen produces a worse transcript than the same candidate on a headset in a quiet room, regardless of which platform or tool does the transcribing. If accuracy on a video interview looks poor, check the audio setup before you blame the software.
Phone screens: no platform to lean on
A phone screen has no meeting platform doing the recording for you, which is both the complication and the opportunity. The complication: you need a separate method to capture the audio, whether that's your phone's own call recording feature where available, a business phone system's recording function, or capturing the audio on your computer if you're calling through a softphone or a headset plugged into your laptop. The opportunity: because nothing is capturing it by default, you choose exactly how and where the audio and transcript live, rather than inheriting a meeting platform's retention settings. For the specific methods and what each one announces to the candidate, see how to record a phone screen.
Once you have an audio file from a phone screen — a call recording, a voice memo, whatever format your method produces — transcribing it after the fact works the same way regardless of source: upload the audio file to a transcription tool or service and it returns text, usually with timestamps and, if the tool supports it, separate labels for each speaker.
In-person interviews: the one nobody's platform records
An onsite interview has no call to record, no meeting platform, and no default transcript of any kind — which is likely why it is the interview stage recruiters take the thinnest notes on. Two practical approaches:
- Record audio directly, then transcribe. A phone set to record, or a laptop with recording software open, placed between the interviewer and candidate, captures the conversation the same way a phone screen recording would. Transcribe the resulting file afterward through the same tools you'd use for a phone screen recording.
- Capture it live on a device in the room. Some transcription apps run live from a laptop's own microphone, producing text while the conversation happens rather than after. This works for a one-on-one in a quiet office; it degrades quickly in an open floor plan or a loud panel room with several people talking over each other.
Whichever method you choose, tell the candidate before you start. An unannounced phone on the table reads very differently from one you mentioned in the invitation or said out loud when they sat down. For wording, see our interview recording consent script.
What determines accuracy, regardless of method
Four things move transcription accuracy more than which tool you pick:
- Microphone distance. A phone or laptop mic more than a few feet from the speaker picks up room echo along with the voice, and echo is what most automatic transcription struggles with.
- One voice at a time. Overlapping speech — a panel interrupting each other, a candidate and a parent in the next room — is the single most common cause of garbled sections in an otherwise clean transcript.
- Uncommon words. Names, certifications, product names and acronyms specific to the role are the words automatic transcription gets wrong most often, even in an otherwise accurate transcript. Skim these before you quote one in a scorecard or submittal.
- Background noise you control. Notifications, a second monitor with a video playing, a door left open to a hallway — these are all fixable in the two minutes before a call starts, and fixing them does more for accuracy than switching tools.
When automatic transcription is not accurate enough
A human transcription service is the third method, and it earns its cost in a specific situation: a heavy accent on both sides, a candidate on a poor phone line, or an interview you plan to quote extensively in a legal or compliance context where an automated tool's occasional misheard word is a real risk rather than a minor annoyance. Human transcription services typically charge per minute of audio and return a file within a stated turnaround window, often same-day to a few days depending on the service and rush option chosen. For a routine phone screen or video interview, automatic transcription is usually accurate enough that paying per minute for a human transcriber is not worth the wait; for the interview you expect to revisit in a dispute, it can be worth it.
A middle option worth knowing about: several automatic tools let you correct and save specific terms — a candidate's name, a certification, a product name your team uses constantly — so future transcripts in the same account get them right. If your team interviews for the same handful of roles repeatedly, teaching the tool your field's vocabulary once pays off across every interview afterward, and costs nothing extra on most plans that offer it.
File formats, and what to do with them afterward
Transcription tools commonly output one of a few formats: plain text, VTT or SRT (caption formats with timestamps, common from video platforms), or a structured export with speaker labels built in. For an interview, timestamps and speaker labels matter more than which specific format you get, because they are what let you jump back to the exact moment a candidate said something before you quote it.
Once you have a transcript, three things are worth doing before you file it away:
- Skim it once for the uncommon-word errors described above, especially anything you plan to quote.
- Pull the two or three lines that actually support your scorecard scores, rather than keeping the whole transcript as your only record of evidence.
- Decide how long you're keeping it and where, consistent with how long you keep interview notes generally — see how long to keep interview notes.
A transcript that nobody reads again is not worth the storage or the privacy exposure. The point of transcribing is to make specific moments findable later, not to accumulate a file you never open.
Where a live question guide changes the workflow
If you're running a structured interview from a written question bank, live transcription that is tied to that question guide — ticking off each question as the transcript shows it was asked and answered — removes a specific failure mode: realizing after the call that you skipped a must-ask question because you were focused on listening rather than checking a printed list. This is the scenario Interview Signal is built around: it builds the question guide from the job description, ticks off coverage during the call from the live transcript, and turns the transcript into a scorecard with quotes checked against it, all without a bot joining the meeting. That is one specific use of live transcription, not a requirement for transcribing an interview at all — the methods above work whether or not you use any particular tool.
Questions people ask
Do I need special equipment to transcribe an in-person interview?
No. A phone propped between you and the candidate, or a laptop with a decent built-in microphone, is enough for a one-on-one conversation in a quiet room. The audio setup matters more than the device: two people talking at a normal distance from one microphone, with the door closed and notifications off.
How accurate is automatic transcription for a candidate interview?
For a clear recording of two English speakers with no heavy accents or crosstalk, general-purpose speech-to-text is usually accurate enough to read without correction, though names, certifications and product names are the most common errors and are worth a quick pass before you rely on a quote.
Should I transcribe every interview, including phone screens?
Transcribe whichever interviews you would otherwise take notes on, since the transcript replaces the notes, not just supplements them. A 15-minute phone screen benefits from a transcript as much as an hour-long onsite, arguably more, because it is the stage where recruiters take the thinnest notes.
Can I transcribe a recording after the interview instead of live?
Yes, and it is often more accurate, since after-the-fact transcription is not competing with your attention on the conversation. The trade-off is that you cannot see must-ask questions ticked off in real time, which live transcription with a question guide gives you.