Sound Design and Why It Carries Half the Emotion
What sound design actually consists of, why it is the first thing cut and the last thing that should be, and how it makes generated footage feel physically real.
Sound is the half of a film that audiences experience without noticing, which is precisely why it is the first thing removed when budgets tighten. A viewer can articulate that a shot looked good. They cannot usually articulate that a scene felt real because someone built a room tone, placed footsteps, added a distant air conditioner and put a subtle low frequency bed under the wide shot. They simply believe the scene, and belief is what the film was for.
A finished commercial soundtrack is made of four layers that are constructed separately and combined. Dialogue or voiceover carries the explicit message. Music carries the emotional arc and sets the pace. Sound effects give individual events physical presence. Ambience, the continuous background of a place, gives the scene a location and, more importantly, makes the absence of sound feel like a room rather than like a technical fault. Removing any one of these is immediately audible even to people who cannot name what is missing.
For generated and composited footage, sound does something additional and specific. Generated imagery has no recorded audio at all, so every sound in the finished film is a deliberate placement. This is an opportunity rather than a burden: a shot of a machine can be given exactly the mechanical character the brand wants, and a shot of a product being handled can be given a weight and a materiality that sells the object. Correctly designed sound is one of the most effective ways to make generated footage stop reading as synthetic, because the ear supplies the physical reality the eye is uncertain about.
The evidence on the commercial value of audio is stronger than most budget conversations assume. Zhang et al. (2025), analysing audiovisual features of short video advertising on TikTok, found that audio characteristics contribute measurably to consumer engagement behaviours. Xiao et al. (2026) reached compatible conclusions studying user engagement from a combined visual and audio perspective on short form platforms. These are not findings about production polish, they are findings about whether people engage.
Music selection should follow the edit's structure rather than the client's taste, and this is a frequent point of friction. The useful criteria are tempo against the cut rate, whether the track has a usable build and a resolution point, whether it leaves midrange space for the voiceover, and whether it ends rather than fades. A track that a client loves but which has no build cannot support a film that needs to arrive somewhere, and no amount of editing rescues it.
Licensing is a legal matter that gets treated casually and occasionally becomes expensive. A track needs a licence covering the actual use: the channels, the territories, the duration and whether the use is paid media. Library music with a proper commercial licence is the normal solution. Music taken from a streaming service is not licensed for commercial video regardless of how the film is distributed, and a client who supplies a track they like should be asked where it came from before it is cut into anything.
Voiceover is the most audience facing audio decision and deserves genuine casting. Peng et al. (2025) found that vocal cues influence dynamic credibility judgements, which is the research statement of something producers observe constantly: the same script read by two performers produces two different levels of trust. Direction matters as much as casting. A read that is too polished sounds like advertising and triggers scepticism; a read that is slightly conversational tends to be believed.
The mix is where all of this either works or collapses, and its main job is priority. At any moment one element should be dominant and the others supporting, and the mix moves that priority around as the film progresses. The commonest amateur mistake is a mix in which music and voiceover compete at similar levels throughout, which makes the narration tiring to follow and the music emotionally inert.
Loudness standards matter more than they used to because the same film plays across contexts with different requirements. A mix that is correct for a cinema is too quiet for social, and a mix optimised for a phone speaker sounds thin and harsh in a ballroom. The practical answer is to master separately for the main destinations rather than exporting one file everywhere, and to check the event version on a large system before the event rather than after.
The argument for protecting the sound budget is straightforward. Colour, graphics and additional shots all improve a film incrementally. Sound determines whether the film feels real, and a film that does not feel real is not persuading anyone regardless of how it looks. When a budget has to be cut, cutting a shot is almost always less damaging than cutting the mix.
References
Zhang, Z., Qiu, K., & Ye, Y. (2025). Influence of audiovisual features of short video advertising on consumer engagement behaviors: Evidence from TikTok. Journal of Business Research, 201, Article 115662. https://doi.org/10.1016/j.jbusres.2025.115662
Xiao, L., Li, X., & Mou, J. (2026). Exploring user engagement behavior with short-form video advertising on short-form video platforms: A visual-audio perspective. Internet Research, 36(1), 154–188. https://doi.org/10.1108/INTR-07-2023-0521
Peng, Z., Wang, C., & Jiang, X. (2025). On how vocal cues impact dynamic credibility judgments: Mouse-tracking paradigm examining speaker confidence and gender through voice morphing. Journal of Speech, Language, and Hearing Research, 68(11), 5261–5277. https://doi.org/10.1044/2025_JSLHR-24-00849