Voiceover Direction and Casting for Corporate Video
How to cast a voice for a corporate film, what to say in the session to get a usable read, and the technical decisions that make narration sound expensive.
The voice is the most exposed decision in a corporate film. Everything else can be adjusted, cut around or graded, but the narration runs continuously and the audience forms a judgement about the company from it within a sentence. Despite this, voice casting is often handled as a procurement task rather than a creative one, with a shortlist chosen from demo reels and a session booked without direction.
The first decision is what the voice is meant to establish, and it is a positioning question rather than a preference. Authority, warmth, energy, calm competence and approachability are different products and they call for different voices. A technical B2B film usually needs measured credibility rather than enthusiasm. A consumer brand may need the opposite. Deciding this before listening to demos prevents the common outcome, which is choosing whichever voice sounded most pleasant in isolation.
Demo reels are useful for eliminating and unreliable for choosing, because they are edited compilations of a performer's best work in contexts unlike yours. The reliable method is a paid custom read of thirty seconds of the actual script from a shortlist of three or four. This costs a small amount, takes a day, and reveals immediately which voice suits the material. It also reveals which performer takes direction well, which matters more over a full session than tonal quality does.
Vocal characteristics carry meaning that operates below conscious evaluation. Peng et al. (2025) found that vocal cues shape dynamic credibility judgements, examining how confidence conveyed in the voice influences how listeners assess a speaker. The practical version of this is that pace, pitch stability and the handling of sentence endings all affect whether narration is believed. A read that rises at the end of statements sounds tentative. A read that is uniformly emphatic sounds like advertising and triggers resistance.
The direction that produces the best results is almost always about the listener rather than about the performance. Asking for a warmer read produces a self conscious result. Asking the performer to speak to one specific person, explaining something they find genuinely useful, produces a natural one. Similarly, giving the reason behind a line, telling the performer what the client actually wants the audience to feel at that moment, works better than adjectives about tone.
Pace should be established against picture rather than in the abstract. The common error is recording narration at a comfortable reading pace and then discovering it does not fit the edit, which forces either a rushed re-record or a re-cut. Recording against a rough edit, or at minimum with the intended durations marked in the script, means the read arrives with the right rhythm. It also lets the performer pause where the film needs a pause rather than where the punctuation suggests one.
Multiple takes should be planned rather than treated as failure. A standard approach is a full read for consistency, a second full read with different energy, and then targeted pickups on the lines that matter most, particularly the opening, the key proposition and the call to action. Editors assemble the final narration from across takes, and having genuine alternatives for the important lines is what makes that possible.
For multilingual projects, casting per language is not optional if quality matters. A performer reading a language they do not speak natively will produce something recognisably wrong to the audience it was made for. Recording all languages in the same week against the same picture keeps the versions consistent in pace and energy, which is what makes them feel like one film rather than three unrelated ones.
Technical decisions determine whether the voice sounds expensive. A treated room rather than an untreated one, a microphone suited to the voice rather than whichever is available, consistent distance and level throughout, and a mix that gives narration priority over music are what separate professional narration from competent narration. Zhang et al. (2025) found that audiovisual features contribute measurably to engagement, and narration quality is one of the audio features an audience registers without identifying.
The alternative worth considering honestly is synthetic narration, which has improved substantially and is genuinely adequate for internal training, temporary scratch tracks and high volume low stakes content. For customer facing brand communication it remains a compromise, both because the performance lacks the specificity a director can obtain and because Kirk and Givi (2025) documented that perceptions of AI authorship can shape consumer responses to marketing communications. Using it for a scratch track during editing and casting a real voice for the delivered film is the practical position for most corporate work.
References
Peng, Z., Wang, C., & Jiang, X. (2025). On how vocal cues impact dynamic credibility judgments: Mouse-tracking paradigm examining speaker confidence and gender through voice morphing. Journal of Speech, Language, and Hearing Research, 68(11), 5261–5277. https://doi.org/10.1044/2025_JSLHR-24-00849
Zhang, Z., Qiu, K., & Ye, Y. (2025). Influence of audiovisual features of short video advertising on consumer engagement behaviors: Evidence from TikTok. Journal of Business Research, 201, Article 115662. https://doi.org/10.1016/j.jbusres.2025.115662
Kirk, C. P., & Givi, J. (2025). The AI-authorship effect: Understanding authenticity, moral disgust, and consumer responses to AI-generated marketing communications. Journal of Business Research, 186, Article 114984. https://doi.org/10.1016/j.jbusres.2024.114984