LinkTranscriptLinkTranscript
← All guides

Language Learning with a Transcript and a Second Caption Track

September 9, 2026 · 7 min read

Textbook dialogues are slow and nobody talks like them. YouTube has a thousand hours of people talking like people in any language you might be learning, and a good share of those videos carry two caption tracks: one in the language spoken, one in a translation. That pair is a bilingual reader with audio attached. Here is how to use it.

Picking a video

You need a video with an uploaded caption track in the target language and a second track in a language you read well. YouTube's caption menu (the gear icon on the player, then Subtitles) lists every track the video has; that is where to check before you start. LinkTranscript shows one track per video, the one the video serves by default, which for most videos is the language being spoken. The language chip beside the word count tells you which one you got.

Two things to check. First, that the target-language track is uploaded rather than automatic. Automatic captions are what speech recognition heard, so for a learner they teach the recognizer's mistakes. The auto-generated versus uploaded guide explains how to tell them apart; the short version is that punctuation and capital letters mean a person was involved. Second, that the translation track is a real translation. YouTube's player can auto-translate any track on the fly, and those machine translations are listed under "Auto-translate" in the caption menu; a track listed by language name alone was uploaded by the creator, which is usually far better.

Good sources: creators who caption their own work for an international audience, language-teaching channels, news programs that publish scripts, and talks with professional subtitles. Length matters less than density; five minutes of clear speech about one topic beats an hour of banter.

The chips row at the top of the example transcript: word count, reading time, and the caption language
The language chip shows which track you are reading. The other tracks a video carries are in YouTube's own caption menu.

The two-column method

Start in the target language. Paste the link on LinkTranscript, leave timestamps on, and read the transcript without the video playing. Do not stop at every unknown word. Read a full paragraph, decide what you think it means, and mark the two or three words that blocked you. The copy control on each line (it appears when you hover) takes the line with its timestamp, so you can paste your marked lines into a note as you go.

Then check against the translation. On the YouTube page, open the transcript panel (the "...more" link under the title, then "Show transcript"); when a video has several tracks, a language menu at the bottom of that panel switches between them. Timestamps are the alignment: a line at 2:14 in one track corresponds to 2:14 in the other, give or take a second. Find your marked lines and check your guess against the translation. Being wrong here is the useful part; the words you guessed wrong are the ones you will remember.

Which track to read first depends on where you are. A beginner can flip the order: read the translation first to know what the passage is about, then read the target language knowing the meaning and let the words attach to it. An intermediate learner should start in the target language as above, because the guessing is the exercise. An advanced learner can skip the translation entirely, listen first with no text, then read the target-language transcript to catch what the ear missed.

If you prefer both columns on screen at once, export the target-language transcript as plain text with timestamps on (Export, then TXT) and open it beside the YouTube transcript panel set to the translation. Lines line up by time.

Shadowing with the player

Reading gives you vocabulary. Shadowing gives you the sounds. Pick one paragraph you now fully understand, click the timestamp on its first line, and let the player run while you read along in the target language, half a second behind the speaker, matching the rhythm and the stops. Click the timestamp again and repeat. Four or five passes on a single paragraph does more for pronunciation than an hour of passive listening, because you are producing the sounds with a model playing in your ear.

The timestamp click is what makes this practical. Scrubbing a player back to the start of a sentence by hand takes longer than the sentence.

Turning the unknowns into a list

By the end of a session you have a note full of lines with timestamps and marked words. Turn it into vocabulary before you close the tab, while the context is fresh.

For each word: the word, the sentence it appeared in (already copied), the translated version of that sentence from the other track, and the timestamp. Keep the sentence. You will remember "the one about the ferry schedule" long after you would remember the word alone. If you use Anki or another flashcard app, the front is the target-language sentence with the word blanked out, the back is the full sentence and the translation. The Markdown export (Export, then MD) gives you the whole transcript with the source link and language at the top if you want the full text in your notes app.

A twenty-minute session

Five minutes: read one section in the target language, mark the blockers.

Five minutes: open the translated track on YouTube, check the marked lines, correct your guesses.

Five minutes: shadow the paragraph you understood best, four passes.

Five minutes: build the cards from today's marked lines.

That is one section of one video. Tomorrow, the next section, same video. A ten-minute video lasts a week this way, and by the end you will know it cold.

What the transcript will not do

It will not translate for you, and it shows one track per video; the other tracks live in YouTube's caption menu. It will not tell you which track is better; that is the punctuation test above. And it does not carry the pictures, so a cooking video where the speaker points and says "this one" is a poor choice for the reading steps and a fine one for shadowing. Pick videos where the words carry the meaning, and the method works from the first session.

Try it on a video

Paste a YouTube link and get a clean, exportable transcript in seconds.