Narration

Have your document narrated, and keep the mp3.

Studio-grade synthesis on our servers, returned as a file you can download, attach, or publish. Your browser's built-in speech can play audio but cannot save it.

Why this is a server-side tool

Browsers ship a speech synthesiser, and it is genuinely useful for skimming something yourself. It has two limits that matter the moment the audio is for anyone else: the voices are robotic, and the Web Speech API exposes no way to capture the output. No in-browser tool can hand you a file — that is the API, not a paywall.

This tool synthesises on our servers with the same voices our rendered videos use, and returns an mp3. That is a real cost per run, which is why it needs an account.

How long documents are handled

One synthesis covers about 5,000 characters, which is roughly twelve minutes of narration. Longer documents are split into sections at paragraph boundaries, so a section never begins mid-sentence, and you narrate them one at a time.

The tool shows how many sections your document became before you start, rather than silently narrating the first chunk and stopping.

  • Female and male narrator voices.
  • Listen time shown per section before you synthesise.
  • The exact text that will be read is displayed alongside.

Audio or video?

Audio is the right format for documents you need to get through yourself — commuting, walking, anywhere your eyes are busy. It is cheap to produce and needs no visuals.

It falls apart on anything visual. A narrator saying 'as Figure 3 shows' to someone with no Figure 3 has communicated nothing, and spoken figures do not stick. If your document's finding lives in a chart, the video version is the one that works.

Frequently Asked Questions

Can I download the audio as an MP3?

Yes — that is the point of this tool. Each section comes back as an mp3 with a download button.

How long can the document be?

Any length. One synthesis covers about 5,000 characters, so longer documents are split into sections at paragraph boundaries and narrated one at a time.

Are these better than my computer's built-in voices?

Substantially. These are the same synthesised narrator voices used in our rendered videos. Device voices flatten questions and mispronounce technical terms.

Does it work on a scanned PDF?

No — a scan has no text layer, so there is nothing to read aloud. Run OCR over it first.

llms.txt