Voice Over Scripts: Writing for the Ear Rather Than the Page
How spoken script differs from written copy, the constructions that fail when read aloud, and how to mark up a script for a performer.
A voiceover script is not copy that happens to be read aloud. It is written for a listener who cannot re-read, cannot control the pace, and is simultaneously watching something. Copy written for the page and handed to a performer produces a read that is technically correct and difficult to follow, and the fault is in the writing rather than the delivery.
The first difference is sentence length and construction. A listener holds a sentence in working memory until it resolves, so a long sentence with subordinate clauses is a demand rather than a flourish. Short sentences, one idea each, active voice, and the subject near the verb are not stylistic preferences here; they are what makes the sentence parseable in real time.
The second is that the listener has no punctuation. A comma and a full stop are indistinguishable to the ear except through the performer's phrasing, which means a sentence whose meaning depends on punctuation will be ambiguous when spoken. Rewriting so the structure is carried by word order rather than by marks removes the ambiguity.
The third is that numbers and abbreviations have to be written as they will be said. A figure like 1,250,000 will be read differently by different performers, and an abbreviation may be spelled out or spoken as a word. Writing one point two five million, and marking whether an initialism is spelled or said, prevents a re-record over something the script never specified.
Reading aloud is the only reliable test and it should happen before the client approves rather than in the booth. It reveals the phrases that are awkward to say, the alliteration that becomes noticeable, the sentence that runs out of air, and the real duration rather than an estimate. A script that has never been spoken has not been checked.
Timing should be verified with a stopwatch against the intended runtime. A comfortable commercial read runs roughly one hundred and forty to one hundred and sixty words per minute, and a film that also has to show something needs pauses, which brings a sixty second film to around one hundred and twenty words. Approving a script without timing it is approving a runtime nobody has confirmed.
Markup is the part that most improves a session and is usually absent. Marking the intended emphasis, the pauses, the pronunciation of any name or technical term, and the intent of each section, gives the performer information the words do not carry. Peng et al. (2025) found that vocal cues shape dynamic credibility judgements, and the cues a performer produces depend on understanding what each line is doing.
Direction in the session should be about the listener rather than about the performance. Asking for a warmer read produces self consciousness; asking the performer to explain this to one person who is interested produces warmth. Similarly, giving the reason behind a line, telling them what the audience should feel at that moment, works better than adjectives about tone.
For multilingual delivery the script should be written to be transcreated rather than translated, and the writer should know that at the outset. Idiom, wordplay and rhythm do not survive literal conversion, and a line that depends on any of them will be flat in the other languages. Writing the primary language slightly short also matters, because the same content occupies different durations and an edit cut to the frame will not accommodate the others.
The final check before recording is to read the script against the storyboard and remove every sentence that describes what the picture already shows. Narration duplicating the image is the most common form of length in a video script, and cutting it is the fastest route to a film that fits its runtime without losing anything the audience was receiving.
References
Peng, Z., Wang, C., & Jiang, X. (2025). On how vocal cues impact dynamic credibility judgments: Mouse-tracking paradigm examining speaker confidence and gender through voice morphing. Journal of Speech, Language, and Hearing Research, 68(11), 5261–5277. https://doi.org/10.1044/2025_JSLHR-24-00849