Upscaling and Frame Interpolation in AI Video

What upscaling and frame interpolation can genuinely recover, where they introduce artefacts, and how to use them without making footage look processed.

Generated video frequently arrives at a lower resolution and a lower frame rate than the delivery specification requires, and the finishing stage has to close that gap. Upscaling and frame interpolation are the two tools for this, they work well within limits, and understanding those limits prevents the characteristic over processed look that makes enhanced footage recognisable.

Upscaling increases pixel dimensions. Traditional methods interpolate between existing pixels and produce a softer version of the same image. Model based upscalers do something different: they infer plausible detail that was not present in the source, reconstructing texture, edges and fine structure based on what they have learned. This is genuinely useful and it is also a form of invention, which is where the risks sit.

The invention becomes a problem on specific content. Faces are the clearest case, because an upscaler reconstructing facial detail can subtly change a person's appearance, and on a recognisable individual that is unacceptable. Text is the second, since letterforms may be reconstructed into shapes that are almost but not quite correct. Fine repeating patterns, fabric weave, mesh, brickwork, can acquire a regularity they did not have. The rule is to upscale environments and textures freely, and to be cautious with faces, text and anything the audience will scrutinise.

The realistic limit for quality is roughly a doubling of linear dimensions. Going from a smaller generated resolution to full high definition, or from high definition toward four thousand pixels wide, generally holds up. Attempting a fourfold increase produces a plasticky, over sharpened result with visible haloing around edges, and it is usually better to deliver at a lower resolution than to deliver something that has obviously been stretched.

Frame interpolation increases frame rate by generating intermediate frames between existing ones. It is used to bring generated footage from a lower rate up to the twenty four, twenty five or thirty frames the delivery format requires, and to create slow motion from material that was not shot for it. On smooth, predictable movement it works well. On complex motion it fails in a recognisable way: limbs and edges warp, objects that pass in front of each other tear, and fast motion produces a rubbery quality.

The failure modes concentrate where the algorithm cannot determine what moved where. Occlusion, meaning one object passing in front of another, is the hardest case because the newly revealed area has to be invented. Fast rotation, water, particles, smoke, and any scene with many independently moving elements are all difficult. Shots with a single subject moving predictably against a stable background interpolate almost invisibly.

There is an aesthetic consideration beyond artefacts. Very high frame rates produce the hyperreal quality that audiences associate with broadcast video rather than with film, and applying heavy interpolation to a brand film can make expensive material feel cheap. The frame rate is a stylistic decision, and matching the destination convention matters more than maximising smoothness.

Order of operations in the finishing chain affects the result. Upscaling and interpolation should happen before grading and before any grain or texture is applied, because these processes analyse the image and are confused by added noise. Applying them after a grade also risks amplifying whatever the grade emphasised. The reliable sequence is to enhance the raw material, then assemble, then grade the finished timeline as a unit.

The consistency requirement is the same one that governs everything else in generative production. Enhancement applied unevenly across a film produces shots that do not match, and mismatched sharpness is as visible as mismatched colour. Kirk and Givi (2025) found that perceptions of authenticity shape consumer responses to AI generated marketing communications, and inconsistent processing is one of the signals that invites exactly that scrutiny. Whatever settings are used should be applied uniformly and checked across the whole sequence.

The most useful practical guidance is to generate as close to the delivery specification as the tools allow rather than planning to fix it afterwards. Enhancement is a recovery technique, not a production strategy, and every stage of recovery costs some quality. A shot generated at the right size and rate needs nothing, and a shot generated small and stretched twice will look it, regardless of how good the enhancement tools have become.

References

Kirk, C. P., & Givi, J. (2025). The AI-authorship effect: Understanding authenticity, moral disgust, and consumer responses to AI-generated marketing communications. Journal of Business Research, 186, Article 114984. https://doi.org/10.1016/j.jbusres.2024.114984