The Real Limits of AI Video and When You Still Need a Live Shoot
An honest list of what generative video still cannot do reliably, and the specific situations where a camera remains the only correct answer.
A studio that will not name the limits of its own tools is selling rather than advising. Generative video has become genuinely capable, and there remain categories of work where it is the wrong instrument, not because the output looks poor but because the thing being asked for is not what the technology produces. Knowing those categories saves clients money and prevents films that fail for reasons nobody anticipated.
The first limit is documentary truth. A camera records that something happened. A generative model produces something plausible. When the purpose of a shot is to serve as evidence, that a factory operates, that a team exists, that a customer uses the product, that an event took place, generation cannot do the job, however convincing the result. This is not a quality threshold that will be crossed by better models; it is a difference in what the two things are.
The second limit is identifiable real people. Staff, founders, customers and spokespeople must be themselves, both because audiences can verify and because synthesising a person carries legal and ethical exposure. The U.S. Copyright Office (2024) addressed the digital replicas question directly in its report on that subject. Kirk and Givi (2025) found that perceptions of AI authorship shape consumer responses to marketing communications and can produce negative reactions in some conditions, which for testimonial content is a commercial risk on top of the legal one.
The third limit is exact brand geometry and legible text. Logos, typography, packaging copy and product proportions are precise structures, and a probabilistic process approximates them. This is solved by compositing rather than by better prompting, and any studio promising generated shots with perfectly rendered brand marks is promising something the technology does not reliably do.
The fourth limit is sustained human performance. A person delivering a line with genuine emotional continuity over thirty seconds, reacting authentically, holding a gaze, is beyond what generated material sustains without drift. Short generated shots of people work; a performance does not. Films built around a spoken message from a human being should be filmed.
The fifth limit is verified physical behaviour. How a product actually deploys, how a mechanism actually moves, how a material actually falls, how a liquid actually pours. Generation produces plausible physics, and plausible is a problem when a technical audience knows the real behaviour. For engineered products this is where 3D built to specification outperforms both a camera and generation, because it is accurate and controllable.
The sixth limit is regulatory contexts where claims must be substantiated. Medical, financial, safety and food categories require that what is shown corresponds to what is approved, and a reviewer will ask where a visual came from. Generated material in these contexts is not prohibited, but it must be clearly illustrative and reviewed accordingly, and it cannot depict outcomes, procedures or results.
Against those limits sit the things generation does better than any alternative, and they are substantial. Environments that would require travel, permits or construction. Products that do not physically exist yet. Concepts, futures and abstractions. Visual exploration before commitment, where ten directions can be tested for less than the cost of scouting one location. Connective and atmospheric material that would otherwise consume a shoot day. Nguyen et al. (2026), in a cross national study of responses to AI generated advertising, found that viewer response depends heavily on execution and perceived authenticity rather than on the technology itself, which is the empirical version of the point: used for the right shots and executed properly, audiences respond to the work rather than the method.
The practical consequence is that most strong projects are hybrid, and the hybrid should be designed rather than arrived at. A typical structure films the people, the premises and the proof, builds the product and the mechanism in 3D, generates the worlds and the atmosphere, and then unifies everything through one grade and one sound mix so the joins disappear. Audiences almost never identify which shots came from where, which is the correct outcome.
The decision framework is a shot by shot test at storyboard stage. Does this shot need to be true. Does it contain an identifiable real person. Must any text or brand mark be exact. Does it depend on verified physical behaviour. Is it subject to regulatory review. A yes to any of those routes the shot to a camera or to specified 3D. Everything else is a candidate for generation. Making that call forty times produces a film that is both affordable and defensible; making it once for the whole project produces one or the other.
References
U.S. Copyright Office. (2024). Copyright and artificial intelligence, Part 1: Digital replicas. https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-1-Digital-Replicas-Report.pdf
Kirk, C. P., & Givi, J. (2025). The AI-authorship effect: Understanding authenticity, moral disgust, and consumer responses to AI-generated marketing communications. Journal of Business Research, 186, Article 114984. https://doi.org/10.1016/j.jbusres.2024.114984
Nguyen, K. M., Phan, T. M., Tran, Y. N. N., Nguyen, A. T., Nguyen, T. L. N., Hoang, G. H., Tran, T. T., & Nguyen, N. T. (2026). Evaluating the efficacy of AI-generated advertising: A cross-national analysis of customer responses on brand perceptions and customer engagement with evidence from Vietnam and Australia. Journal of Global Scholars of Marketing Science, 36(2), 293–341. https://doi.org/10.1080/21639159.2026.2617659