Motion Graphics, Lower Thirds and On Screen Text
How on screen text should behave, the sizing and timing rules that make it readable, and why graphics are where brand consistency is enforced.
On screen text is where an otherwise strong film most often looks cheap. The footage may be well shot and well graded, and then a lower third appears in a default typeface with a default animation and the whole thing drops a tier. Graphics are the element the audience reads rather than merely sees, and reading invites judgement.
The first rule is that text must be on screen long enough to be read twice. A viewer needs to notice the text, read it, and confirm it, and the common failure is text timed to the editor's reading speed on a large monitor rather than to a viewer's on a phone. The practical minimum for a short name and title is around three seconds, and longer strings need proportionally more.
Size should be set for the smallest destination rather than the largest. A lower third that is comfortable in a 16:9 master becomes marginal in a vertical crop and unreadable on a phone in daylight. Designing at the size the audience will actually see, and checking on an actual device rather than a desktop preview, catches this in a minute.
Typography should come from the brand guideline and frequently does not, because the brand book specifies print typefaces without considering motion. Thin weights that read beautifully on paper disappear against moving footage and disintegrate on an LED wall. The practical resolution is to define the video typography once, as part of the brand system, specifying the weights that survive motion and compression rather than the ones that look best on a page.
Animation should be consistent across every title in a film, and consistency is more important than sophistication. One entrance behaviour, one exit behaviour, one duration, applied everywhere. Films where each title animates differently look assembled rather than designed, and the inconsistency is visible even to viewers who could not describe it. This is also what makes graphics reusable across a content library rather than rebuilt per project.
Placement has to respect the destinations. Platform interface elements occupy the lower portion of the frame on social, captions may occupy the same region, and event screens may have physical obstructions. Keeping graphics within a defined safe area, and testing against a real platform preview rather than a clean player, avoids the common outcome where the name of the person speaking sits behind a play button.
Contrast against moving footage is the technical problem that plain text cannot solve, because the background changes underneath it. The reliable solutions are a subtle scrim behind the text, a consistent placement over an area of the frame that has been kept simple during shooting or generation, or a graphic container that belongs to the brand system. Relying on a drop shadow to rescue white text over bright footage produces something that is legible and looks like an afterthought.
Text is also where accuracy obligations concentrate, particularly in regulated categories. Names, titles, figures, claims, legal lines and disclaimers all appear as text and all get checked. Building a check of every on screen string against a source document, before the final render, is a five minute task that prevents the most embarrassing category of error, which is a misspelled name of someone senior.
The relationship to comprehension is worth taking seriously rather than treating graphics as decoration. Zheng et al. (2022) found that adding subtitles to audio visual material assists comprehension, and Beege and Ploetzner (2025) examined how design and cognitive load affect what viewers take from video. Text that duplicates the narration word for word competes with it, while text that supplies what the narration cannot, a figure, a name, a term, supports it. Deciding which role each piece of text plays is a design decision with a measurable effect.
For multilingual delivery, graphics determine how expensive the language versions are. Text baked into a shot cannot be replaced without regenerating or re-rendering it; text on a separate layer, in a container designed with room to expand, becomes a swap. Since string lengths differ substantially between English, Bahasa Malaysia and Mandarin, designing the containers for the longest expected string rather than the first one written is what keeps the additional versions cheap.
References
Zheng, Y., Ye, X., & Hsiao, J. H. (2022). Does adding video and subtitles to an audio lesson facilitate its comprehension? Learning and Instruction, 77, Article 101542. https://doi.org/10.1016/j.learninstruc.2021.101542
Beege, M., & Ploetzner, R. (2025). Learning from interactive video: The influence of self-explanations, navigation, and cognitive load. Instructional Science, 53(1), 99–119. https://doi.org/10.1007/s11251-024-09693-5