· The AI Transcript team
How YouTube captions actually work (and why that matters for transcripts)
Every transcript tool is really a caption tool. Here is what YouTube publishes, the difference between a written and an automatic track, and what that means for the text you get back.
Every "AI transcript" tool for YouTube is doing one of two things. Either it is listening to the audio and writing text itself, or it is reading text YouTube already published. The second is far more common, because it is instant and free. This site does the second, and understanding what that text is explains almost everything about the result you get.
What a caption track actually is
When you turn subtitles on in the YouTube player, the words are not burned into the picture. They arrive as a separate timed text file — a list of cues, each with a start time, an end time and a fragment of speech:
start 1.20s end 3.36s "All right, so here we are"
start 5.31s end 7.97s "in front of the elephants"
That file is what a transcript tool reads. The timestamps in your transcript are not estimated or aligned by anyone; they are the platform's own, copied across. That is why clicking one lands exactly on the right moment.
Two kinds of track, and the gap between them
A video can carry either kind, or both.
Creator-written captions are typed or corrected by a human. They have punctuation, capitalisation, correct spellings of names, and speaker changes where it matters. Transcripts from these are usually good enough to quote with a spot check.
Automatic captions come from YouTube's speech recognition. They are impressive and they are also consistently weak in four specific places: proper nouns, numbers, technical jargon, and any moment where two people talk at once. They also arrive with almost no punctuation, which is why an automatic transcript reads as one long breathless run.
The practical difference is not "one is 95% accurate and the other 90%". It is that the errors cluster exactly on the words you are most likely to want to quote. A tool that does not tell you which kind you are reading is hiding the one fact you need.
Why some videos return nothing at all
Three common cases, none of which a better tool can fix:
- The video is too new. Automatic captions are generated after upload and can take minutes or hours. There is nothing to read until they land.
- The creator disabled captions. Some do. The track simply is not published.
- The captions are burned into the picture. Common on short-form video. Those words are pixels, not text, and only optical character recognition or speech recognition could recover them.
What normalisation has to fix
Raw caption data is not ready to read. Automatic tracks scroll — a line appears, then reappears a moment later with more words attached — so the same phrase can occur two or three times. Cues sometimes carry a duration of zero. Occasionally an end time lands before its own start.
Anything presenting this as a transcript has to collapse the duplicates, repair the broken timings from the neighbouring cues, and collapse the whitespace, without ever touching the words themselves. Get that wrong and you either show the reader stuttering repeats or you quietly edit what someone said.
What this means when you use one
- Check the label. If the transcript is marked auto-generated, treat names and numbers as suspect.
- Verify quotes against the audio. A timestamp click takes ten seconds and catches the errors that matter.
- If a video returns nothing, it is almost always the video, not the tool. Try a second video before concluding anything.