
Accessibility
Captions and accessibility for document video
Text-heavy video has accessibility problems that talking-head video does not, and most of them are invisible to the person who made it.
The problem specific to document video
A talking-head video with captions is broadly accessible: the captions carry the audio, and the visual channel is a person. Document video is different, because the visual channel carries information — a chart, a number, a diagram — that captions do not describe.
So a viewer using a screen reader, or watching without being able to see the frame clearly, gets the narration and misses the evidence. This is the failure that automated accessibility checks do not catch.
Fix one: narrate the visuals
The cheapest and most effective fix. Instead of 'this chart shows the trend', say 'costs rose steadily from January, then jumped fourteen per cent in September'. The narration now contains the information the chart contains.
This is better writing regardless of accessibility — a viewer looking at their phone one-handed on a train also benefits — which is why it is the fix to do first.
Fix two: captions that are actually correct
Auto-generated captions are around 90–95% accurate on clean audio and considerably worse on technical vocabulary, which is what document video is made of. A caption track that renders your product name three different ways is worse than useless for search and for readers.
If the video was generated from a script, the exact text exists — upload it rather than auto-generating. That is a thirty-second step that removes the entire problem.
- Upload a script-derived SRT rather than auto-generating.
- Two lines maximum, around seven words each.
- One second minimum per cue.
- Position away from the bottom edge, where platform UI overlays sit.
Fix three: contrast and type size
Document video inherits document typography, and document typography assumes a page at reading distance. WCAG AA wants 4.5:1 contrast for normal text — grey-on-white body copy from a report frequently fails, and looks fine to the person who made it because they know what it says.
Type size is the related problem. Anything intended to be read on a phone needs to be roughly twice its slide size. If you cannot fit the text at that size, the frame has too much text.
Fix four: pacing
Text on screen needs to be readable at the slowest reasonable reading speed, not yours. The working rule is about 180 words per minute of on-screen text — noticeably slower than average silent reading — plus a beat before the scene changes.
A frame that holds three lines therefore needs at least four or five seconds, regardless of how long the narration takes. When narration is shorter than that, extend the hold rather than cutting early.
The transcript, again
Publishing the full transcript on the page is the single accessibility measure that also does the most for SEO, and it serves people who cannot use video at all for whatever reason — bandwidth, environment, preference.
If the narration describes the visuals properly, as above, the transcript is a genuinely complete alternative to the video rather than a partial one.
Frequently Asked Questions
Are captions legally required?
In many contexts yes — public sector bodies, education, and large employers in several jurisdictions. Requirements vary, but the practical answer for anything published is to caption it.
Are auto-captions good enough?
Not for technical content. Accuracy drops sharply on domain vocabulary and proper nouns. If you have the script, uploading it is faster than correcting auto-captions anyway.
How do I make charts accessible in video?
Narrate what the chart shows in specific terms — the direction, the magnitude, the turning point — rather than referring to it. That single change covers most of the gap.