Skip to content

Transcript cleaner no timestamps, no labels, in paragraphs.

Paste a transcript or open an SRT, VTT, SBV, TXT or JSON file. The cleaner strips timestamps (00:01:02, [01:02]), cue numbers and arrows, speaker labels and tags like [Music] and (laughs), then joins broken lines into paragraphs. You see the before and after counts right away, and copy or download the result.

Free, no sign-up
What to clean
Before
-
After
-
How to use it

How to clean up a transcript

  1. 01

    Paste the text or open the file

    A transcript copied from YouTube, an SRT, VTT or SBV subtitle, a TXT with timestamps or a transcript JSON all work. The format is detected for you.

  2. 02

    Pick what to remove

    Timestamps, speaker labels, sound tags and hesitations each have a switch. Choose paragraphs, one sentence per line or the original lines.

  3. 03

    Copy or download

    The clean text updates with every option. Compare the word count before and after, copy it in one click or download a TXT.

How do I remove timestamps from a YouTube transcript?

When you copy a transcript from a YouTube video, every chunk comes with its timestamp on a line of its own (0:00, 0:04, 1:15) and the text is broken into short pieces. Sometimes there is also a line like "1 minute, 5 seconds", the text meant for screen readers. Paste it all here: the timestamps and those lines go, and the text becomes paragraphs.

Timestamps are recognized in every common shape: [00:01:02], (1:02), 00:01:02,500, 01:00:00:12 with frames, ranges such as 00:01 - 00:05, and the SRT and VTT timing lines with the --> arrow. A time inside a sentence, like "the show starts at 10:30", stays in the text because it is not a transcript timestamp.

If you need the timestamps, turn the first option off: each paragraph then starts with the time where it begins, as [01:02], or [01:02:03] once the recording passes the one hour mark.

In the transcriptIn the clean text
0:04 to today's sessionto today's session
[00:01:02] Good morningGood morning
00:00:01,000 --> 00:00:03,000(removed)
the show starts at 10:30the show starts at 10:30

How do I remove speaker labels and [Music] tags?

Meeting, interview and podcast transcripts usually open each turn with a label: "Speaker 1:", "SPEAKER_00:", "Interviewer:" or the person's name. You can remove every label, remove only the generic ones (and keep real names) or keep them all. Even when a label goes, a change of speaker still starts a new paragraph, so the conversation does not blur together.

To avoid deleting real text, a name only counts as a label when it is generic, when it opens more than one turn or when most lines carry one. A single "Note: bring a dish" inside ordinary text stays where it is.

Sound tags go too: anything in square brackets, like [Music] and [Applause], parentheses that describe a sound, like (laughs), (coughs) and (inaudible), and song lyrics between ♪ marks. Parentheses with real content, like (2019) or (see chapter 3), stay.

How do I join broken subtitle lines into paragraphs?

Subtitles are made for the screen: each cue holds two short lines and the sentence carries on in the next cue. To read, edit or publish it, that has to become running text. The paragraph option joins every line and starts a new paragraph every few sentences (four by default, you choose). A blank line in the original and a change of speaker always start a paragraph.

Auto-generated captions often come with no punctuation at all. Then there are no sentences to count, so paragraphs are cut by length, about twenty words for each sentence you asked for. You can also ask for one sentence per line, handy for reviewing or building a script, or keep the lines as they are and only strip the clutter.

The last option handles the finish: it removes double spaces and spaces before commas and periods, drops the comma left behind when a hesitation goes, and capitalizes the start of each sentence (and a lone "i" in English). Abbreviations such as Dr. and p. are not mistaken for the end of a sentence.

Can I remove "um" and "uh" from a transcript?

Yes, with the hesitation option, which starts off. It removes only sounds that carry no meaning in any sentence: in English, um, uh, erm and hmm; in Portuguese, ahn, hum and é... (with the dots, since é alone is a verb); in Spanish, eh, mmm and este... with the dots. The language is detected from the text itself.

Words that are sometimes a filler and sometimes not, like "like", "so", "you know" and "actually", stay. Deleting them all would change sentences such as "I like this light". To see how often they show up, use the filler word counter.

Frequently asked questions

How do I copy a YouTube transcript without timestamps?

Open the video's transcript on YouTube, select all the text and copy it. Paste it here with timestamp removal on: the timestamps, the screen reader lines and the breaks disappear, and the text becomes paragraphs you can copy or download as TXT.

How do I turn an SRT file into plain text?

Open the SRT file or paste its content. Cue numbers, the --> timing lines and italic tags are removed, and the lines of each cue are joined into sentences and paragraphs. It works the same with VTT, SBV and transcript JSON files.

Does cleaning change the words?

No. It only removes what you choose: timestamps, speaker labels, sound tags and, if you turn it on, hesitations such as um and uh. The finishing option touches spaces, leftover commas and capital letters at sentence starts. No word is replaced or corrected.

Can I keep the speaker names?

Yes. Choose to keep the labels, or remove only the generic ones such as Speaker 1 and SPEAKER_00 and keep real names. Either way, every change of speaker starts a new paragraph.

Why did my text end up as one big paragraph?

Most likely the text had no punctuation, which is common in auto-generated captions. Without periods, paragraphs are cut by length. Lower the number of sentences per paragraph or keep the lines to see the original breaks.

Is the transcript cleaner free?

Yes. It is free, with no sign-up and no usage limit. A FilmScribe plan only makes sense if you want us to transcribe the recording, with who spoke and the time of every sentence.

Have the video? The transcript comes out clean. With who spoke and when.

Upload the file or paste a YouTube link. We transcribe word for word, separate the speakers, split it into chapters, and you download a TXT with or without timestamps and names.

See the result on your own recording before adding a card.