VTT to TXT converter just the speech, no header, no timings.
To convert VTT to TXT, paste the contents of the .vtt file or drop the file in and click convert. We remove the WEBVTT line, NOTE and STYLE blocks, timings, position settings and voice tags, and give you only the speech in UTF-8, with or without a timestamp at the start of each line.
Free, no sign-upSecond file
For example, the video name and the version up for approval.
This text has no timings. We will build subtitles from reading speed: each line or sentence becomes a block of up to 2 lines.
The JSON has a time for every word. We cut blocks of up to 2 lines of 42 characters, at sentence ends and pauses.
Subtitle report
- Subtitles
- Duration
- Words
- Average CPS
- Max CPS
- Flagged
| # | Time | Duration | Text | Characters per line | CPS | Flags |
|---|
How to convert VTT to TXT
- 01
Bring the .vtt
Drop in the file you downloaded from YouTube Studio, Vimeo, a recorded meeting or a course platform, or paste its contents. Timings without hours, like 01:12.480, are read too.
- 02
Pick lines or paragraphs
One line per subtitle, with [00:01:12] in front if you want, or paragraphs that break at every pause of 2 seconds or more.
- 03
Copy or download the .txt
The text comes out in UTF-8, ready for Word or Google Docs. The report shows how many subtitles and words the file had.
What's inside a VTT file besides the text?
WebVTT was built for web players, so it carries more than an SRT does. At the top sits the WEBVTT line, sometimes followed by metadata such as Kind: captions and Language: en. Between subtitles you may find NOTE blocks with reviewer comments, STYLE blocks with CSS rules and REGION blocks that define areas of the screen.
Each subtitle can also have an identifier on the line above it and settings after the timing, such as line:85% or align:start, telling the player where to draw the text. None of that is speech. When converting, we read only the blocks that have a timing, skip the rest and join each subtitle's lines into one sentence.
lesson-07.vtt
WEBVTT
Kind: captions
Language: en
NOTE reviewed on March 12
lesson-07-01
00:01:12.480 --> 00:01:15.020 line:85% align:center
<v Teacher>Do you remember last week's lesson?</v>
00:01:15.300 --> 00:01:18.000
Open your notes
to page 42.
lesson-07.txt
Do you remember last week's lesson?
Open your notes to page 42. Does the speaker's name stay in the TXT?
It depends on how the name was written in the VTT. Some systems mark the speaker with a voice tag, like <v Ana Souza>, which is format markup rather than subtitle text. Transcripts downloaded from Microsoft Teams usually look like this. When converting, the whole tag goes away, name included, and the TXT keeps only the speech.
Others write the name into the text itself, as in Ana Souza: good morning. That is the usual layout in transcripts of Zoom recordings. There the name is plain text like any other and stays in the TXT. One thing to watch: the option to remove sound descriptions also removes uppercase names followed by a colon, like ANA:, at the start of a line. Leave it off if you want to keep those names.
| In the VTT | In the TXT |
|---|---|
| <v Ana>Good morning, everyone.</v> | Good morning, everyone. |
| Ana: Good morning, everyone. | Ana: Good morning, everyone. |
| ANA: Good morning, with sound removal on | Good morning |
| <i>Opening music</i> | Opening music |
| & and < | & and < |
Where the VTT files people turn into text come from
Almost always from something that was streamed or recorded on the web. VTT is the format that site players, course platforms and meeting tools use to store captions and transcripts, which is why so many people end up with a .vtt when what they wanted was text to read.
- Captions downloaded from YouTube Studio or Vimeo, to become a blog post, summary or description
- A recorded meeting's transcript, to write up the minutes or the list of decisions
- Lecture captions from a course platform, to study with the text alongside
- A webinar or live stream with captions, to pull quotes and audience questions
- Captions of a corporate video, to review the wording with the client in a document
Timings without hours, broken accents and repeated lines
In VTT the hour is optional: 01:12.480 and 00:01:12.480 are the same moment. We read both, and the timestamps in the TXT always come out in full, as hours, minutes and seconds. The encoding is detected for you too: UTF-8, UTF-16 or, when the file is not valid UTF-8, Windows-1252. The text comes out in UTF-8, so accented names stay intact.
Some rolling auto-captions repeat each line across two consecutive blocks to create the scroll-up effect. The converter returns what is in the file and does not merge repeated lines, so those repeats show up in the TXT. If that's your case, delete the duplicates while reviewing, or start from the recording itself, which we transcribe without repeats.
More than pulling the text out of a VTT. The transcript comes straight from the recording.
See also
Frequently asked questions
How do I convert VTT to TXT?
Paste the contents of the .vtt or drop the file in, leave TXT as the output and click convert. The text shows in the preview so you can copy it, and the button downloads a UTF-8 .txt. It's free and needs no sign-up.
Do the WEBVTT line and NOTE blocks end up in the text?
No. The header, metadata like Kind and Language, NOTE, STYLE and REGION blocks, subtitle identifiers and position settings are all left out. The TXT gets only what would be shown on screen as speech.
Why did the speaker's name disappear from my TXT?
Because your VTT kept the name inside a voice tag, like <v Ana>, which is markup and goes away in the conversion. If the name had been written in the text, as in Ana: good morning, it would stay. To get speakers separated from the recording itself, FilmScribe transcription labels who spoke in each passage.
Can I keep timestamps in the TXT?
Yes. Turn on the timestamp option and each line starts with the subtitle's start time, as [00:01:12], without milliseconds. With paragraphs on, the time appears only at the start of each paragraph.
My TXT has repeated lines. Why?
Your VTT most likely repeats lines on purpose, as some rolling auto-captions do. The converter doesn't merge repeats, so it never deletes a line that was really said twice. Remove the duplicates while you review.
Does it work with Zoom or Teams VTT transcripts?
Yes. Both deliver transcripts as timed .vtt files, and the converter reads them like any other WebVTT file. The difference is the speaker's name: written in the line, it stays; inside a voice tag, it goes.
Have the recording, not just the VTT? We transcribe it with speakers.
Upload the video or audio, or paste a YouTube link. We transcribe every line with who said it and when, and you download TXT, PDF or SRT.
See the result on your own recording before adding a card.