Putting the same words in both channels is the most common mistake in document video. The split that actually works.

Craft

What goes on the slide, and what goes in the voiceover

Viewers read at 240 words a minute and listen at 150. Put the same text in both places and they finish reading first, then leave.

The arithmetic behind the rule

Silent reading runs at roughly 240 words per minute. Narration runs at 150. Put a forty-word paragraph on screen and read it aloud, and the viewer finishes it in ten seconds while the narrator takes sixteen. For six seconds they have nothing to do, and this repeats on every slide.

That dead time is what makes auto-generated slideshows feel lifeless. It is not the template or the voice — it is that one of the two channels is always idle.

The split that works

Give each channel a different job. The screen holds the claim: short, declarative, scannable. The narration holds the reasoning: the why, the caveat, the example.

The viewer reads the claim in two seconds, then spends the remaining twelve listening to why it is true while the claim stays visible as an anchor. Both channels are working, and neither is repeating the other.

  • On screen: the claim. Under fifteen words per line, two or three lines maximum.
  • In narration: the reasoning, the evidence, the qualification.
  • Overlap: only the key term or the number, deliberately repeated for retention.
  • Never: the same sentence in both.

The exceptions

Numbers should appear in both. A spoken figure does not stick, so saying 'thirty-four per cent' while '34%' is on screen is reinforcement rather than redundancy.

Technical terms and proper nouns too. A viewer who has only read a term needs to see the spelling while hearing the pronunciation, and vice versa.

Direct quotations should be on screen and not narrated, or narrated and not on screen — reading a quote aloud while it is displayed is the standard redundancy again, and quotes are long enough for it to hurt.

How this changes what you write

It means you are writing two things from one source, not adapting one thing. Take a section of your document and produce a three-word headline, two condensed lines, and a full narration paragraph. They come from the same prose and none of them is the prose.

The mechanical part of this — finding the strongest sentences for the on-screen lines while keeping the full text as narration — is what our storyboard tool does, and seeing the split laid out per slide is usually more convincing than the argument for it.

Testing it

Watch your video on mute. If it still makes sense, you have put too much on screen and the narration is decorative. Then listen with your eyes closed. If that also makes complete sense, the screen is decorative.

A well-split video fails both tests slightly, and that is the point: each channel needs the other.

Frequently Asked Questions

How many words should be on a video slide?

A headline plus two or three lines under about fifteen words each. Around forty words total is the ceiling for something read comfortably on a phone.

Should captions count as on-screen text?

They are a separate layer and do duplicate the narration by design, for accessibility and mute viewing. Keep them visually distinct from your content text so they read as captions.

What if my content is genuinely text-heavy?

Then it needs more slides, not fuller ones. A dense section becomes four scenes with one claim each rather than one scene with a paragraph.