A/B Testing Video Creative Properly

How to run a video creative test that produces a real answer: one variable, enough delivery, the right metric and an honest read of the result.

Most video creative tests produce a number rather than an answer. Two assets run, one performs better, a conclusion is drawn, and the difference was within the range that random variation would produce anyway. Testing badly is worse than not testing, because it generates confident beliefs that are wrong and then those beliefs shape the next year of production.

The first requirement is a single variable. If version A has a different opening, a different length, a different call to action and different music, then a difference in performance identifies nothing, because four things changed. A real test changes one element and holds everything else identical. This is unglamorous and it is the entire basis on which the result means something.

The highest value variable to test first is the opening, because it moves performance more than anything else and is cheap to vary. Four different openings against one identical body, judged on early retention, will typically produce a spread large enough to be meaningful and actionable. The next most useful variables are the format and aspect ratio, the call to action, and the audio bed, roughly in that order.

The second requirement is enough delivery behind each variant to distinguish signal from noise. A test where each version received a few hundred impressions has produced nothing regardless of how different the percentages look. The practical approach for most brands is to keep the number of variants small so that each receives meaningful volume, rather than splitting a modest budget across eight versions and learning nothing about any of them.

The third requirement is the right metric. Judging creative on final conversions conflates the creative with targeting, landing page, offer and price, and requires far more data than most campaigns have. Judging on early retention isolates the creative and produces usable answers quickly. Kim et al. (2025), studying playback interactions in mobile video viewing, found that behaviour during playback carries information that aggregate view counts obscure, which supports reading the retention curve rather than a single headline figure.

Placement has to be held constant or the test is measuring the placement. Frade et al. (2023) found that in stream ad format and placement materially affect visual attention and effectiveness, and Davtyan et al. (2025) documented differences between skippable, non skippable and brand placement strategies. Running version A as pre roll and version B in a feed and comparing them is not a creative test, it is a placement comparison with the creative confounded.

Simultaneous testing beats sequential testing for a reason that is easy to overlook. Running A this week and B next week introduces every difference between those weeks: seasonality, competitor activity, news, platform changes, audience saturation. Running both at once, split randomly across the same audience, removes all of it. Sequential testing is sometimes unavoidable and its results should be held loosely.

Audience fatigue is a real effect that can invert a result if a test runs long. Yin et al. (2023) found that skippable advertising influences advertising avoidance intention, meaning that repeated exposure carries a cost that accumulates. A test run long enough for one variant to reach heavy frequency in a small audience is measuring fatigue rather than creative quality, which is an argument for shorter tests at controlled frequency.

The result should be recorded in a form that survives the campaign. A one line note saying which variable was tested, what the variants were, what the difference was and what the team concluded is what turns individual tests into accumulated knowledge. Brands that do this for a year know what their particular audience responds to. Brands that run tests and report them in a slide deck learn the same thing repeatedly.

Finally, some things should not be tested because the answer is already known or the test would be dishonest. Whether captions help is settled. Whether a hook matters is settled. Whether a claim can be made is a legal question rather than a performance one. Reserving testing for genuine uncertainty, and acting on established findings without re-litigating them, is what makes a testing programme efficient rather than a permanent state of doubt.

References

Kim, E., Oh, S., & Park, S. (2025). An empirical study of user playback interactions and engagement in mobile video viewing. IEEE Access, 13, 78272–78289. https://doi.org/10.1109/ACCESS.2025.3566402

Frade, J. L. H., Oliveira, J. H. C. de, & Giraldi, J. de M. E. (2023). Skippable or non-skippable? Pre-roll or mid-roll? Visual attention and effectiveness of in-stream ads. International Journal of Advertising, 42(8), 1242–1266. https://doi.org/10.1080/02650487.2022.2153529

Davtyan, D., Tashchian, A., & Thomas, M. L. (2025). A comparative analysis of skippable ads, non-skippable ads, and brand placements: Evaluating YouTube advertising strategies. Journal of Advertising Research, 65(3), 464–478. https://doi.org/10.1080/00218499.2025.2464276

Yin, S., Li, B., & Zhou, Q. (2023). The impact of skippable advertising on advertising avoidance intention in China. Marketing Intelligence & Planning, 41(8), 1121–1137. https://doi.org/10.1108/MIP-07-2022-0298