Why Human Creative Direction Still Decides AI Video Quality
Where the quality difference between two studios using identical tools actually comes from, and what direction consists of in a generative workflow.
Two studios with the same models, the same budget and the same brief will produce work of very different quality, and the difference is not access to tools. It sits in a set of judgements that happen before, during and after generation, none of which the model makes. Understanding what those judgements are is useful for a client choosing a supplier and for anyone trying to work out why their own output is inconsistent.
The first judgement is what the film should be. A model can produce a beautiful shot of almost anything, which is precisely the problem: without a decision about what this particular audience needs to understand and feel, the output is attractive material that solves no commercial problem. Vakratsas and Wang (2021), writing on artificial intelligence in advertising creativity, framed AI as operating within the creative process rather than replacing the judgement that governs it, and that framing has held as the models have improved.
The second judgement is the visual language. Left to itself a model produces a generically pleasing average of everything it associates with the words it was given. A direction that specifies the light, the palette, the lens character, the material vocabulary and what is excluded produces work that looks like a particular studio made it for a particular brand. This is the decision that most visibly separates professional output from prompt output, and it is made as still images before any motion exists.
The third judgement is knowing when a shot is right. A generation session produces many attempts, most of which are acceptable and a few of which serve the storyboard. Selecting well requires knowing what the shot is for, which is information the model does not have. Studios that use whatever arrives are optimising for effort; studios that reject most of what they generate are optimising for the film.
The fourth judgement is structural: what goes where, how long each shot holds, where the audience's attention should be at each moment. Yu et al. (2025) found that visual attention to advertising messages within video stories is distributed unevenly rather than remaining constant, which means the ordering of information is a decision with measurable consequences. No amount of shot quality compensates for putting the important thing where nobody is looking.
The fifth judgement is knowing which problems to solve with which tool. Nour (2026), comparing prompt engineering with model selection, found that these are distinct levers with different effects rather than interchangeable ones. The practical version is that some failures are structural to the approach and no amount of rewording fixes them, and recognising that quickly is the difference between an afternoon of progress and an afternoon of frustration.
The sixth judgement is where accuracy is non negotiable. Brand marks, product geometry, packaging copy, real people and anything functioning as evidence must be exact or filmed, and deciding this shot by shot is a direction decision rather than a technical one. Kirk and Givi (2025) found that perceptions of authenticity shape consumer responses to AI generated marketing communications, which is why the studio's judgement about where to draw that line has commercial consequences rather than merely aesthetic ones.
The seventh judgement is finishing, and it is the least discussed. Forty independently generated shots do not belong together until someone grades them as a unit, applies consistent grain, matches the lens character, and builds a sound design that gives the material physical presence. A film that looks coherent does so because of decisions taken after generation finished.
The commercial reading of this for a client is that the useful diagnostic question when commissioning is not which model a studio uses. Models change every few months and everyone has access to broadly the same ones. The question is what happens between the brief and the first frame, how the visual language is established and recorded, how shots are selected and rejected, and who is accountable for the decisions in between.
Nguyen et al. (2026), in a cross national study of responses to AI generated advertising, found that viewer responses depend heavily on execution and perceived authenticity rather than on the technology itself. That is the whole argument in one finding. The tools removed the budget excuse for poor work; they did not remove the requirement to be good, and the requirement to be good is met by people making decisions.
References
Vakratsas, D., & Wang, X. S. (2021). Artificial intelligence in advertising creativity. Journal of Advertising, 50(1), 39–51. https://doi.org/10.1080/00913367.2020.1843090
Yu, W.-Y., Wang, Z. J., & Tao, C.-C. (2025). The dynamics of visual attention to advertising messages in video stories. Journal of Advertising, 54(5), 713–731. https://doi.org/10.1080/00913367.2025.2524837
Nour, R. R. (2026). Prompt engineering versus model selection for cognitive accessibility in large language models: An empirical study. IEEE Access, 14, 44740–44754. https://doi.org/10.1109/ACCESS.2026.3667133
Kirk, C. P., & Givi, J. (2025). The AI-authorship effect: Understanding authenticity, moral disgust, and consumer responses to AI-generated marketing communications. Journal of Business Research, 186, Article 114984. https://doi.org/10.1016/j.jbusres.2024.114984
Nguyen, K. M., Phan, T. M., Tran, Y. N. N., Nguyen, A. T., Nguyen, T. L. N., Hoang, G. H., Tran, T. T., & Nguyen, N. T. (2026). Evaluating the efficacy of AI-generated advertising: A cross-national analysis of customer responses on brand perceptions and customer engagement with evidence from Vietnam and Australia. Journal of Global Scholars of Marketing Science, 36(2), 293–341. https://doi.org/10.1080/21639159.2026.2617659