How to transcribe an interview: format, example and APA rules
To transcribe an interview, upload the recording to FilmScribe. We return the text word by word, with every turn attributed to a speaker and timestamped. You review it against the audio, swap speaker labels for IDs (I, P1), apply your notation and export TXT or PDF. Then de-identify the text and format quotes and appendices following APA 7.
Step by step
- 01 Upload the recording In FilmScribe, upload the interview audio or video, named by participant ID, and select just the range you need.
- 02 Get text with speakers We transcribe word by word, with a timestamp on every word and every turn attributed to a speaker.
- 03 Review against the audio Click a line to hear it, fix names and terms in the text and mark [inaudible] wherever you can't be sure.
- 04 Swap speakers for IDs Rename speakers to I, P1 and P2 (in the app on Studio and Agency) and keep the timestamp on every turn.
- 05 Export, de-identify and format Download TXT, apply your notation, remove identifying details and format quotes and appendices following APA 7.
In qualitative research, the transcript is your data: your codes, themes and quotes all come from it. In FilmScribe, you upload the recording and get the text back word by word, with every turn attributed to a speaker and timestamped. Typing drops out of the picture, and your time goes to what reviewers actually judge: checking, labeling, de-identifying and formatting.
This guide covers that workflow and everything after it: verbatim or clean, a full example, notation, APA 7 rules for quotes and appendices, and ethics.
In short
- Upload the recording to FilmScribe and get text with speakers separated and a timestamp on every turn.
- Choose verbatim or clean verbatim before you start. Switching halfway means redoing everything.
- Use consistent speaker IDs (I, P1, P2) and keep the timestamp at the start of every turn.
- In APA 7, participant quotes from your own study are not personal communications and don’t go in the reference list.
- Quotes of 40 words or more become block quotes. Full transcripts, if included, go in an appendix.
- De-identify everything and follow what your consent form promised. The review against the audio is still yours.
How to transcribe an interview with FilmScribe
1. Upload the recording
Upload the audio or video in formats like MP3, WAV, M4A, MP4 and MOV, or paste a YouTube link. Name the file by participant ID and date, like P4_2026-08-14.wav, never by real name: it shows up in your file list. If there is small talk before and after the interview, select just the part you need. Only that range is processed and counts toward your hours.
2. Get text with speakers
We transcribe word by word, with a timestamp on every word, and attribute each turn to a speaker. The language is detected on its own, among 90+ languages, which helps with bilingual fieldwork. See transcription and speakers.
3. Review against the audio
Read while the recording plays. Click any line and the player jumps to it. Fix proper names, acronyms and field-specific jargon right in the transcript, and search for every mention of a term. When you can’t be sure, write [inaudible] with the timestamp instead of guessing. Star the lines you already know you’ll quote.
4. Swap speakers for IDs
On Studio and Agency, rename “Person 1” to “I” and “Person 2” to “P3” in the app, or ask us to identify speakers by name. On Pro, do the swap in your text editor after exporting. Filter the transcript by speaker to read one participant’s answers on their own.
5. Export
Download TXT to apply your notation and formatting in your word processor, or PDF to send to your advisor. Need the audio of one quote for a conference talk? Download just that snippet as MP3.
6. Store it carefully
Files are kept for 60 days. That gives you room to review, but download what you need and make sure your storage matches your consent form.
Want to try it on one of your interviews? The free audio to text page leads to the trial of 3 days with 1 hour of material, with a card, and you see the result on your own interview.
What a good interview transcript does
Reviewers look for three things: the transcript is consistent (same rules across every interview), traceable (any quote can be found in the audio) and ethical (no participant can be identified without consent). In journalism the emphasis shifts: the transcript protects you when a source says they were misquoted, and timestamps plus an unedited original are your evidence.
Answer these in your methods notes:
- Verbatim or clean verbatim, and why?
- What notation will you use for pauses, laughter, crosstalk and inaudible audio?
- How are participants labeled and de-identified?
- Who transcribed, and how was the transcript checked? If you used automatic transcription, say which and how you reviewed it.
Verbatim vs clean verbatim
| Aspect | Verbatim | Clean verbatim |
|---|---|---|
| Keeps | Everything: “um”, “like”, “you know”, false starts, repetitions, pauses, laughter | The content, in readable sentences |
| Removes | Nothing | Fillers, stutters, false starts |
| Best for | Discourse and conversation analysis, linguistics, some clinical research | Thematic analysis, case studies, market research, journalism |
| Review time | Longer | Shorter |
| Risk | Hard to read and quote | Losing nuance that was data (a long pause before a sensitive topic) |
A common compromise: keep the transcript verbatim, then lightly clean the passages you quote in the paper, marking omissions with ellipses. If you do this, say so in your methods.
Clean verbatim is not rewriting. Grammar, slang and dialect stay as spoken. If a quote needs context, add it in square brackets: “she [the principal] never answered.”
Interview transcript example
A semi-structured interview with a public school teacher, for a master’s thesis in education. Verbatim version:
Interview 3. Participant: P3 (teacher, 12 years in the classroom)
Date: 2026-08-14. Length: 00:47:12. Mode: remote (video call)
Transcription: verbatim. Notation: see Appendix A.
[00:00:12] I: Can you start by telling me about your first year at the school?
[00:00:17] P3: So... it was rough, you know. I started in March, the class
already had, like, their teacher [laughs].
[00:00:29] I: Rough in what way?
[00:00:31] P3: In the sense that nobody... nobody explained how the grade
review meetings worked. The coordinator just handed me the
agenda [inaudible 00:00:44] and said "figure it out."
[00:00:47] I: And now?
[00:00:49] P3: [long pause] Now I'm the one explaining it to the new ones.
The same passage in clean verbatim:
[00:00:12] I: Can you start by telling me about your first year at the school?
[00:00:17] P3: It was rough. I started in March, and the class already had
their teacher. [laughs]
[00:00:29] I: Rough in what way?
[00:00:31] P3: Nobody explained how the grade review meetings worked. The
coordinator just handed me the agenda [inaudible] and said
"figure it out."
[00:00:47] I: And now?
[00:00:49] P3: Now I'm the one explaining it to the new ones.
Clean verbatim dropped “you know”, “like” and the repetition without changing meaning. It also lost the long pause before the last answer. If that pause matters to your analysis, verbatim is the right call.
Transcription notation
There is no single official standard for most social science work. Pick a simple set, list it in your appendix and use it everywhere.
| Situation | Common notation | Example |
|---|---|---|
| Unclear audio | [inaudible] plus timestamp | [inaudible 00:12:40] |
| Best guess at a word | Word in parentheses | we drove to (Dayton) |
| Nonverbal sounds | Square brackets | [laughs], [crying], [coughs] |
| Transcriber’s note | Square brackets | [long pause], [points at photo] |
| Interrupted speech | Ellipsis where the sentence breaks | I was going to... |
| Crosstalk | [crosstalk] | [crosstalk 00:22:05] |
| Emphasis | Italics or caps | I did *not* know |
| De-identified detail | Bracketed description | [company name] |
APA style for interview transcripts and quotes
APA 7 draws a clear line between two kinds of interviews. The APA Style guidance on personal communications (checked September 2026) says quotations from research participants in your own study are not personal communications, which is the rule most students get wrong.
Interviews that are part of your research
If the interview is data in your study, quotes are not personal communications. You don’t cite them with a date and you don’t list them in the references. You identify the speaker with a pseudonym or ID and describe the participants in your method section.
- Short quote, in text: P3 said nobody explained “how the grade review meetings worked.”
- 40 words or more: a block quote, on its own line, indented 0.5 inch, no quotation marks.
- Omissions: an ellipsis. Your own clarifications: square brackets.
If you include full transcripts, they go in an appendix after the references, labeled Appendix A, Appendix B and so on, each with a bold title, and referred to by label in the text. The timestamp on every turn lets your committee check any quote against the audio.
Interviews outside a formal study
An interview you did for background, like an email exchange with an expert, is a personal communication. It is cited in the text only, not in the reference list: (J. Rivera, personal communication, March 3, 2026).
Published interviews
An interview published in a magazine, podcast or video is cited like that source (article, podcast episode, video), with a reference list entry.
Ethics and confidentiality
If your study went through an IRB or ethics board, your consent form made promises about how recordings are stored, who can hear them and when they are destroyed. Transcription has to keep those promises:
- De-identify names, nicknames, employers, children’s names, small towns and unique job titles. “The only woman on the city council” identifies someone as well as a name does. Replace them with bracketed descriptions:
[school name],[small town in Ohio]. - Check storage rules before uploading recordings to any online service.
- Keep the key separate from the transcripts.
Interviews recorded on video, a call or YouTube
Many interviews now happen over video calls or on camera. In FilmScribe, the route is the same for all of them:
| Where the interview is | What to do |
|---|---|
| Recorder or phone (audio) | Upload the audio file as is (MP3, WAV, M4A) |
| Camera (video) | Upload the video, or pull just the audio if your connection is slow |
| Recorded video call | Download the recording from your computer or the cloud and upload the file |
| Interview on YouTube | Paste the link; if it’s someone else’s, cite it as a published source |
To pull the audio out of a video first, the free extract audio from video tool gives you an MP3 or WAV with nothing to install. Planning your interview guide? The free words to minutes tool estimates how long your questions take to ask.
Is typing it by hand still worth it?
For one short interview or a few minutes of audio, yes: play it at reduced speed, use a keyboard shortcut to pause and rewind a few seconds, and add a timestamp at every turn. For a dissertation with twenty one-hour interviews, time is what hurts, and text that arrives with speakers separated and timestamps reviews far faster than a transcript typed from scratch. Running a lot of interviews (a research group, a newsroom, oral history)? See interviews.
Common mistakes
- Mixing styles. Two verbatim and three clean transcripts make comparisons shaky.
- Guessing unclear audio.
[inaudible 00:31:10]beats an invented sentence that ends up quoted. - Fixing the participant’s grammar. It changes your data, and sometimes the meaning.
- Dropping timestamps. Without them, nobody can check a quote.
- Real names in the appendix or the file name. Consent to record is not consent to be named.
- Reviewing at the end. Review each interview right after the text comes back, while you still remember the tone.
- Not disclosing software. If you used automatic transcription, say so and describe your review.