Subtitle and Caption Standards for Multi Language Video

The technical standards that make subtitles readable, the difference between burned in and sidecar captions, and the quality checks that catch what automation misses.

Subtitles have shifted from an accessibility feature to a primary delivery format, because a large share of video is watched without sound and because multilingual markets need one film to serve several audiences. Despite that shift they are still frequently produced as a final export setting rather than as a designed deliverable, and the difference is visible to anyone trying to read them.

The technical standards that determine readability are well established and easy to apply. Two lines maximum, because three lines cover too much of the frame and take too long to read. A limit of around forty two characters per line for Latin scripts, fewer for a small screen. A minimum display duration of roughly one second so that lines do not flash, and a maximum of around six or seven seconds so that a line does not sit stale on screen. A reading speed that stays within a comfortable range rather than matching speech exactly when speech is fast.

Line breaks should follow meaning rather than filling the available width. A subtitle broken mid phrase forces the reader to hold an incomplete thought, which costs comprehension even though every word is present. Breaking at natural grammatical boundaries, keeping a subject with its verb where possible, and avoiding a break that leaves a single short word on the second line, are the habits that make subtitles feel effortless.

Position matters more than it used to because platform interface elements occupy predictable regions. Captions placed at the very bottom of the frame will sit behind the description, the username and the control bar on several platforms simultaneously. Raising them into the lower third proper, and keeping them within the horizontal safe area, is what keeps them readable across destinations rather than only in a clean player.

The burned in versus sidecar decision should be made per destination rather than once. Burned in captions are part of the picture: they always appear, they cannot be turned off, they cannot be translated without a re-render, and they are the right choice for social feeds where autoplay is silent and viewers will not enable anything. Sidecar files, such as SRT or VTT, are separate text that the player displays: they can be toggled, they support multiple languages from one video file, they are indexable by search engines, and they are the right choice for websites and platforms that support them. Most projects need both.

The research on comprehension is consistent and supports treating captions as core rather than optional. Zheng et al. (2022) found that adding subtitles to audio visual material assists comprehension, and Pujadas and Muñoz (2020) documented comprehension benefits from captions and subtitles in viewing contexts. For a corporate film delivered to an audience with mixed first languages, which describes most Malaysian workplaces and customer bases, this is a direct argument for budgeting a proper subtitle pass per language.

Typography for multilingual subtitles requires a decision that many projects avoid. Chinese characters need more vertical space and a heavier weight than Latin text to remain legible at the same size, and a typeface chosen for a Latin brand identity may have no Chinese cut at all. Establishing the pairing once at brand level, with defined sizes per script, prevents each project from improvising a different answer and producing versions that do not look related.

Automatic transcription is a starting point and not a deliverable. It reliably mangles product names, technical terminology, company names and any code switching, all of which are routine in Malaysian speech, and it punctuates according to pauses rather than grammar. The workable process is machine transcription followed by human correction against the actual script and a glossary of product and company terms, which is considerably faster than transcribing from scratch and considerably more accurate than accepting the output.

Code switching needs an explicit policy because it occurs naturally in real speech and inconsistent handling is noticeable. When a speaker mixes English into a Malay or Mandarin sentence, the caption can render it as spoken, normalise it to the subtitle language, or mark it. Any of those is defensible and changing approach between scenes is not. Deciding once, before the subtitling begins, produces a set of versions that behave the same way.

The quality check that catches the most problems is to watch the finished video at actual delivery size, on a phone, with sound off, and read it as a viewer would. This reveals lines that are too fast, breaks that read awkwardly, text sitting behind interface elements, and the moments where the captions have drifted out of sync. It takes the length of the film and it is the difference between subtitles that were produced and subtitles that work.

References

Zheng, Y., Ye, X., & Hsiao, J. H. (2022). Does adding video and subtitles to an audio lesson facilitate its comprehension? Learning and Instruction, 77, Article 101542. https://doi.org/10.1016/j.learninstruc.2021.101542

Pujadas, G., & Muñoz, C. (2020). Examining adolescent EFL learners’ TV viewing comprehension through captions and subtitles. Studies in Second Language Acquisition, 42(3), 551–575. https://doi.org/10.1017/S0272263120000042