Captions vs subtitles: what is the actual difference?
Subtitles assume the viewer can hear the audio and only needs the words; captions assume they cannot, so captions also carry who is speaking and the sounds that matter.
People use the two words interchangeably, and most of the time nothing goes wrong. But they do mean different things, and the difference occasionally matters — for accessibility compliance, for how you write the text, and for what a platform expects you to upload.
The short answer
Subtitles assume you can hear the audio and just need the words. They were built for translation: the viewer hears Spanish, reads English.
Captions assume you cannot hear the audio at all. They carry the dialogue and everything else you would need to follow the video — who is speaking, and meaningful non-speech sound.
So: [door slams] is a caption. It is not a subtitle.
What is the difference between closed and open captions?
Cutting across that distinction is whether the text can be turned off.
Closed captions are a separate track the viewer can enable or disable. That is the CC button on YouTube. They come from a subtitle file — SRT or VTT — attached to the video rather than drawn into it.
Open captions are burned into the picture. Every pixel of text is part of the video frame, so there is no way to switch them off, and no player setting can hide them.
This is the distinction that actually changes what you do day to day. Almost all short-form video uses open captions, for one blunt reason: TikTok, Reels and Shorts are largely watched with the sound off, and a viewer scrolling past will never tap a CC button they have not noticed. If the words are not already on screen, they are not being read.
When does the distinction matter?
It matters for accessibility. If you have a legal or institutional obligation — public sector work, education, broadcast — the requirement is usually captions, specifically, and often closed ones. Speaker identification and non-speech audio are not optional extras there; they are the point.
It matters on YouTube. YouTube can read an uploaded subtitle file, which means the words become text associated with the video — searchable and translatable. Burned-in text is pixels, and it is invisible to that. On YouTube the sensible move is both: burn captions in for the muted scroller, and upload an SRT for everything else.
It rarely matters on TikTok, Reels or Shorts. None of them let you upload a subtitle file for a normal post. Burned-in is the only real option, so the terminology question collapses.
What this means for how you write them
If you are producing genuine captions for accessibility, you need more than a transcript: identify speakers when it is ambiguous, and note sound that carries meaning — a phone ringing, laughter, music with lyrics that matter. Ignore incidental noise nobody needs.
If you are producing subtitles for a short-form video, the constraint is different and mostly about legibility. Four to six words on screen at once, sized to read at arm’s length, positioned clear of the platform’s own interface.
Automatic transcription — including what Viten produces — gives you dialogue. That makes it subtitles by the strict definition. For most short-form video that is exactly what you want. If you need true captions for accessibility, treat the automatic output as a first draft and add the speaker labels and sound cues yourself.
Terms worth keeping straight
| Term | Means |
|---|---|
| Subtitles | Dialogue only; assumes the viewer can hear |
| Captions | Dialogue plus speaker IDs and meaningful sound; assumes they cannot |
| Closed | A separate track the viewer can toggle |
| Open / burned-in / hardcoded | Drawn into the picture, cannot be turned off |
| SRT | The most widely accepted subtitle file format |
| VTT | The web format; what HTML5 video requires |
Which should you use?
For short-form social video: burned-in, and call them whatever you like — nobody is checking.
For YouTube: burn them in and upload an SRT. Different jobs, and doing both costs one extra export.
For anything with an accessibility requirement: real closed captions, with speaker identification and non-speech audio, delivered as a file.
Caption your next video with Viten
Automatic captions, AI voiceover and audio clean-up. Free to start, no watermark.