LinkTranscriptLinkTranscript
← All guides

SRT vs VTT: Which Caption File You Need, and How to Fix the Usual Problems

September 8, 2026 · 7 min read

SRT and VTT are the two caption formats almost everything accepts, and the choice between them is mostly about where the captions are going. Both hold the same thing, a list of timed lines of text, so a mistake here is cheap: if an application rejects one, export the other. This guide covers what is inside each file, which tools want which, and the fixes for the problems that come up.

What an SRT file looks like

SubRip Text. A plain text file with a numbered list of cues. Each cue is an index, a time range with a comma before the milliseconds, one or more lines of text, and a blank line. The three lines below are the opening of a NASA video as exported from LinkTranscript.

1
00:00:01,310 --> 00:00:08,700
Between 1968 and 1972, America launched 9 human missions to the Moon, 6 of which successfully

2
00:00:08,700 --> 00:00:12,900
touched down, allowing 12 men to walk on the lunar surface.

3
00:00:12,900 --> 00:00:18,890
NASA's next chapter of lunar exploration, called Artemis, has the task of not just going

SRT dates from the early 2000s and has almost no features beyond the cues, which is exactly why every editor reads it. Some players honor basic HTML-style tags like <i> inside the text; most ignore them.

What a VTT file looks like

WebVTT, the W3C format for captions in HTML5 video. It starts with the line WEBVTT, allows NOTE blocks for comments, uses a period instead of a comma in the times, and does not require the index numbers. LinkTranscript puts the video's title, channel, source link, and generation date in a NOTE block at the top so the file identifies itself.

WEBVTT

NOTE Title: How We Are Going to the Moon - 4K
NOTE Channel: NASA
NOTE Source: https://www.youtube.com/watch?v=_T8cn2J13-4

00:00:01.310 --> 00:00:08.700
Between 1968 and 1972, America launched 9 human missions to the Moon, 6 of which successfully

00:00:08.700 --> 00:00:12.900
touched down, allowing 12 men to walk on the lunar surface.

VTT can also carry positioning and styling (a cue can be told to sit at the top of the frame, or a class can color a speaker's lines), which is why web players prefer it. Those features are optional; a minimal VTT file is an SRT file with a header and periods.

Which one to use where

  • Video editors: SRT. Premiere Pro, DaVinci Resolve, Final Cut Pro (via caption import), CapCut, and Descript all take it. Several also read VTT, but SRT is the one that never surprises you.
  • Websites and HTML5 players: VTT. The <track> element in HTML only accepts WebVTT. Vimeo, Wistia, and JW Player want VTT for captions.
  • Course and learning platforms: usually VTT. Teachable, Kajabi, Thinkific, and Moodle expect it; a few accept SRT too.
  • Desktop and living-room players: either. VLC, Plex, Kodi, and Infuse read both, and will load a subtitle file automatically if it sits beside the video with the same filename.
  • Social platforms that accept caption files: SRT more often than VTT. Check the upload dialog; LinkedIn and YouTube itself both take SRT.
  • People: neither. If someone wants to read the words, send TXT or Markdown, not a subtitle file.

Problem one: the captions are early or late

The captions are timed to the YouTube upload. If the video you are playing was trimmed at the start, or has a title card the upload did not, every cue is off by the same amount. The fix is an offset, not a re-time. VLC has Subtitle > Sub Track Synchronization; Premiere and Resolve let you slide the whole caption track; Subtitle Edit, a free desktop app, has Synchronization > Adjust all times. Measure the gap once on the first line and apply it to all.

Problem two: apostrophes and accents show as junk

The file is UTF-8, which is what every modern player and editor expects. When one asks you to pick an encoding on import, pick UTF-8. If you opened the file in an older editor that saved it as something else, the characters break; re-export from LinkTranscript rather than trying to repair it. Curly quotes and non-English characters are the usual casualties.

Problem three: lines are too long for the screen

Automatic caption cues can run to twenty words. Most players wrap them, but a broadcast or platform delivery spec may require two lines of thirty-two to forty-two characters. Subtitle Edit can re-break cues to a character limit in one pass (Tools > Split long lines), and Premiere and Resolve have a maximum-characters setting on the caption track. Re-breaking changes only the layout; the timings stay attached to the words.

Converting between them

If you have the video link, do not convert; export the format you want from LinkTranscript. If you only have a file, the conversion is mechanical: SRT to VTT means adding the WEBVTT header, swapping the comma in each time for a period, and dropping the index numbers; VTT to SRT is the reverse. Subtitle Edit and ffmpeg ("ffmpeg -i captions.srt captions.vtt") both do it in a second.

The Export row on a LinkTranscript transcript page with TXT, MD, SRT, and VTT buttons
SRT and VTT are one click each in the Export row, with YouTube's own timings.

Where the timings come from

LinkTranscript uses YouTube's own caption timings, so the cues match the video as YouTube plays it. Each cue's end time is its start plus its duration. On the rare track where YouTube omits durations, the tool estimates one from the number of words so no cue has a zero length, which some players reject outright. Uploaded captions are usually timed by a person and break at phrases; automatic ones are timed by the recognizer and sometimes split a sentence across two cues, which is harmless in a player and slightly odd in a transcript.

Try it on a video

Paste a YouTube link and get a clean, exportable transcript in seconds.