JSON to SRT by segment or word by word.
To convert a JSON transcript to SRT, paste the contents or drop the .json file and hit convert. We read each segment's start and end in seconds and, when the file carries word timestamps, build subtitles of up to 2 lines with 42 characters each. The SRT comes back numbered and in UTF-8.
Free, no sign-upSecond file
For example, the video name and the version up for approval.
This text has no timings. We will build subtitles from reading speed: each line or sentence becomes a block of up to 2 lines.
The JSON has a time for every word. We cut blocks of up to 2 lines of 42 characters, at sentence ends and pauses.
Subtitle report
- Subtitles
- Duration
- Words
- Average CPS
- Max CPS
- Flagged
| # | Time | Duration | Text | Characters per line | CPS | Flags |
|---|
How to turn a transcription JSON into an SRT
- 01
Paste or drop the .json
The JSON that AI transcription tools and speech to text APIs return works, with segments holding a start, an end, and text. We recognize the structure on our own.
- 02
Pick segments or words
If the file has word timestamps, check the option to build the subtitles from the words. Without it, each segment becomes one or more subtitles.
- 03
Review the report and download
Check reading speed and long lines for every subtitle, fix overlaps if needed, and download the SRT in UTF-8.
What makes it from the JSON into the SRT
A transcription JSON holds far more than a subtitle file needs. Each segment carries its text, a start and end in seconds (like 4.3 and 6.1), and a pile of technical numbers: token ids, average log probability, the odds that nobody is speaking, compression ratio, and the decoding temperature. An SRT only cares about timing and text.
We turn the seconds into 00:00:04,300, number the subtitles from 1, and trim the leading space that usually opens each text. The technical fields and the detected language are dropped, since SRT has no place to store them.
| In the JSON | In the SRT |
|---|---|
| start and end in seconds | 00:00:04,300 --> 00:00:06,100 |
| text with a leading space | Trimmed text, up to 2 lines |
| words with per-word timing | Subtitles built from the words (optional) |
| tokens, avg_logprob, temperature | Ignored |
| no_speech_prob, compression_ratio | Ignored |
| language | Ignored: SRT has no language field |
Segments or words: which gives better subtitles?
Segments are built for transcripts, not for the screen. One segment can span dozens of seconds of speech and hold a 200 character sentence on a single line. In segment mode, we wrap the text into lines of up to 42 characters, and when it still doesn't fit in 2 lines, we split the segment into several subtitles and share out the time in proportion to the text.
When the JSON has word timestamps, you get better results by building from the words. Each subtitle ends after a period, question mark, or exclamation point, at a pause of 0.7 seconds or longer, or once it reaches 7 seconds, always within 2 lines of 42 characters. The text shows up and leaves the screen right along with the speech.
JSON
{"start": 4.3, "end": 9.8,
"text": " Open your Bibles to John 3:16. Today we will talk about hope."}
SRT
1
00:00:04,300 --> 00:00:06,100
Open your Bibles to John 3:16.
2
00:00:06,400 --> 00:00:09,800
Today we will talk
about hope. Which transcription JSON files we read
We recognize the structure from the content, so the file name and where it came from don't matter. Times in seconds and in milliseconds are each read the right way, and the text encoding (UTF-8, UTF-16, or Windows-1252) is detected as the file loads.
- A segments list where each item has start, end (in seconds), and text: the most common shape from speech to text APIs.
- A top-level words array with a time for every word, next to the segments.
- Words nested inside each segment, when the tool aligns the transcript word by word.
- A transcription list with times in milliseconds under offsets, common in command line tools.
Where to use the SRT you get from the JSON
SRT is the subtitle format almost every app accepts: YouTube Studio uploads it directly, Premiere Pro opens it through File > Import, DaVinci Resolve brings it in from the Media Pool, and CapCut reads it in its captions panel. Many transcription tools can export SRT too, but plenty of people only have the JSON on hand, from an API or a script, or they want subtitles built word by word.
Before you download, take a look at the report. It lists the duration, characters per line, and reading speed of every subtitle, and flags anything over 17 characters per second, lines longer than 42 characters, and subtitles longer than 7 seconds, which are common when a segment ran long. Transcripts sometimes include [Music] or (applause): the option to remove sound descriptions cleans that up in the same pass.
More than a JSON converter. We transcribe the video and hand you the SRT.
See also
Frequently asked questions
How do I convert a JSON transcript to SRT?
Paste the .json contents or drop the file, leave SRT as the output, and click convert. We turn each segment's seconds into 00:00:04,300, number the subtitles, and wrap lines at 42 characters. If the file has word timestamps, check the option to build the subtitles from the words for tighter timing.
Why are my subtitles too long?
Because segments are designed for transcripts and can last dozens of seconds. In segment mode, we split whatever doesn't fit in 2 lines and share out the time by text length. For subtitles that follow the speech closely, export the JSON with word timestamps and use the option to build from the words.
How do I get a JSON with word timestamps?
Turn on word timestamps in the tool or API that made the transcript. The words then show up in a words array, either at the top level or inside each segment. We read both layouts, and the option to build the subtitles from the words becomes available right away.
What happens to avg_logprob, tokens, and no_speech_prob?
They're ignored. Those fields measure how confident the transcription was and help with debugging, but an SRT only holds a number, a time, and text. The detected language is dropped too. If you want to use confidence to review shaky passages, open the original JSON in a code editor before converting.
Can I turn the same JSON into VTT, TXT, or CSV?
Yes. The same JSON can come out as SRT, VTT, TXT, SBV, ASS, TTML, or CSV: just pick a different output format before you download. CSV is handy for reviewing the text in a spreadsheet, and TXT without timestamps becomes plain prose for a blog post, a summary, or a video description.
The transcript has [Music] and (applause). Can I remove them?
Yes. Check the option to remove sound descriptions: we delete anything inside square brackets or parentheses plus lines with ♪, and any subtitle left empty is dropped. In the same conversion you can also strip formatting tags and fix overlapping subtitles.
Skip the JSON: we make the SRT from your video. With speakers and word-level timing.
Upload your video or paste a YouTube link. We transcribe it word by word, in 90+ languages, and you download the SRT, TXT, or PDF without converting anything.
See the result on your own recording before adding a card.