Version Drift: Keeping Output Consistent as Models Update

Why a campaign extended six months later no longer matches, what to record so it can, and how to plan for tools that change underneath you.

A campaign produced in March and extended in September will not match, and the reason is not carelessness. The models have been updated, the defaults have shifted, and the same inputs now produce different outputs. This is a structural property of building on tools that improve continuously, and it needs planning for rather than discovering.

The drift shows up in specific places. Colour rendering shifts. The default aesthetic changes, usually toward whatever the newer version does well. Interpretation of the same description differs. Faces and characters generated from the same reference come out subtly different. Individually these are small; across a set of shots intended to sit beside earlier work they are obvious.

The first defence is to record everything that produced the original work. The reference frames, the conditioning inputs, the settings, the model and version used, and the approved outputs at full quality. A studio with that record can reproduce a matching shot by reusing the structural inputs rather than the description. One without it is starting from a written brief, which is precisely what will be interpreted differently.

The second defence is to lean on conditioning rather than description. A shot generated from an approved reference frame is anchored to something that has not changed, while a shot generated from words is anchored to an interpretation that has. This is the practical reason the stills first workflow matters beyond the initial approval: the approved frames are the campaign's stable reference.

The third defence is to keep the finished assets rather than only the deliverables. High quality masters of every approved shot, not just the edited film, mean that an extension can reuse actual footage from the original campaign rather than regenerating it. For a brand extending a campaign across a year, this is often the simplest answer: recut existing material and generate only what is genuinely new.

Nour (2026), comparing prompt engineering with model selection, found these to be distinct levers with different effects rather than interchangeable ones. Applied to drift, this means that when a new version produces a different result, the fix is usually not to rewrite the description but to change the conditioning or to pin the version, and knowing which lever applies prevents a long unproductive search.

Version pinning, where the tooling allows it, is the cleanest technical answer and is not always available. Where a service offers a specific model version, using it for the duration of a campaign and recording which one was used gives genuine stability. Where only the current version is available, the reference and conditioning approach is the fallback, and it works reasonably well.

The grade is the final unifying layer and it absorbs a surprising amount of drift. Material generated months apart, corrected to a common base and graded as a unit against the same reference frames, will sit together even when the underlying outputs differ. Jonauskaite et al. (2020) documented consistent patterns of emotion associations with colours, which is part of why grading to a documented palette rather than to taste is what makes this reliable rather than lucky.

The credibility argument for taking drift seriously is the same one that applies to consistency generally. Kirk and Givi (2025) found that perceptions of AI authorship shape consumer responses to marketing communications and can produce negative reactions in some conditions, and a campaign whose later assets visibly differ from its earlier ones is exhibiting exactly the signal that invites that scrutiny.

The practical policy for a studio is short. Record the version, settings and reference inputs for every accepted shot. Archive full quality masters of all approved shots, not only the edit. Prefer conditioning over description for anything that may need to be matched later. Pin versions where possible for the duration of a campaign. Grade to a documented palette. Together these turn drift from an unpleasant surprise into a manageable property of the medium.

References

Nour, R. R. (2026). Prompt engineering versus model selection for cognitive accessibility in large language models: An empirical study. IEEE Access, 14, 44740–44754. https://doi.org/10.1109/ACCESS.2026.3667133

Jonauskaite, D., Parraga, C. A., Quiblier, M., & Mohr, C. (2020). Feeling blue or seeing red? Similar patterns of emotion associations with colour patches and colour terms. i-Perception, 11(1), Article 2041669520902484. https://doi.org/10.1177/2041669520902484

Kirk, C. P., & Givi, J. (2025). The AI-authorship effect: Understanding authenticity, moral disgust, and consumer responses to AI-generated marketing communications. Journal of Business Research, 186, Article 114984. https://doi.org/10.1016/j.jbusres.2024.114984