Subtitling and Accessibility for Long Form Recorded Content

What changes when subtitles have to run for forty minutes, the reading speed limits that govern them, and the accessibility obligations attached.

Subtitling a forty minute recorded session is a different problem from subtitling a sixty second film. The volume is twenty times greater, the material is denser and more technical, and the viewer is reading for a sustained period rather than glancing at a line. Standards that are comfortable in short form become fatiguing across long form.

The governing constraint is reading speed, and it is where most long form subtitling fails. Matching speech exactly is possible and produces lines that exceed comfortable reading rates whenever the speaker is fluent, which for a prepared talk is most of the time. Condensing to the sense of what was said, rather than transcribing, is the craft, and it is the difference between subtitles that support the viewer and subtitles that compete with the picture.

The eyetracking evidence is directly applicable to how the lines should be set. Zahedi and Khoshsaligheh (2021) examined the impact of subtitle length and line number on viewers' allocation of visual attention, finding that these parameters affect where viewers look. In a session where the viewer should be looking at a slide or a demonstration, over long subtitle lines actively pull attention away from the thing being explained.

The practical standards follow from that. Two lines maximum. A character limit per line that keeps text readable at the size it will actually be viewed, which for a laptop or a phone is smaller than the editing monitor suggests. A minimum display duration so lines do not flash, and a maximum so a line does not sit stale. Breaks that follow phrase boundaries rather than filling the available width.

Terminology is the accuracy problem specific to technical long form. Product names, chemical names, acronyms, drug names, standards references and company names are exactly what automatic transcription mangles, and they are exactly what the audience is watching for. A glossary supplied before subtitling, and a human correction pass against it, is not optional for this material.

The comprehension benefit is well supported and applies to everyone rather than only to viewers with hearing differences. Zheng et al. (2022) found that adding subtitles to audio visual material assists comprehension, and Pujadas and Muñoz (2020) documented comprehension benefits from captions and subtitles in viewing contexts. For dense technical content delivered by a speaker with an unfamiliar accent, subtitles frequently carry more of the message than the audio.

The delivery format decision differs from short form. Long form recorded content is usually hosted on a platform or a site where a sidecar caption file is supported, which is preferable to burning in: it can be toggled, it supports multiple languages from one video file, and it is indexable by search engines. Burned in captions remain appropriate for the short clips extracted from the session for social.

Publishing the transcript alongside the video is the addition that serves both accessibility and discoverability with one artefact. It gives a viewer who prefers reading a route through the material, it allows searching within the content, and it converts several thousand words of spoken material into indexable text on the page.

Chapters interact with subtitles usefully and are worth building together. A viewer using captions is frequently also navigating rather than watching linearly, and chapter markers let them reach the section they need. Beege and Ploetzner (2025), studying learning from interactive video, examined how navigation and cognitive load influence what viewers take from video material, and navigation is one of the levers that determines retention in long form.

The quality check that catches the most problems is to watch the finished session with the captions on, at actual size, from beginning to end. Reading speed violations, breaks that fall mid phrase, terminology errors, and drift out of sync are all only visible in playback. It takes the length of the session and it is the only reliable verification that the subtitling was produced rather than merely generated.

References

Zahedi, S., & Khoshsaligheh, M. (2021). Eyetracking the impact of subtitle length and line number on viewers' allocation of visual attention. Translation, Cognition & Behavior, 4(2), 331–352. https://doi.org/10.1075/tcb.00058.zah

Zheng, Y., Ye, X., & Hsiao, J. H. (2022). Does adding video and subtitles to an audio lesson facilitate its comprehension? Learning and Instruction, 77, Article 101542. https://doi.org/10.1016/j.learninstruc.2021.101542

Pujadas, G., & Muñoz, C. (2020). Examining adolescent EFL learners’ TV viewing comprehension through captions and subtitles. Studies in Second Language Acquisition, 42(3), 551–575. https://doi.org/10.1017/S0272263120000042

Beege, M., & Ploetzner, R. (2025). Learning from interactive video: The influence of self-explanations, navigation, and cognitive load. Instructional Science, 53(1), 99–119. https://doi.org/10.1007/s11251-024-09693-5