First Frame and Last Frame Transition Techniques
How bounding a generated shot at both ends produces controllable transitions, what it is good for, and the failure modes to plan around.
The most useful control technique in generative video is also the simplest to describe. Instead of asking a model to produce motion from a description alone, you give it the image the shot should start on and the image it should end on, and it generates the movement between them. The shot is bounded at both ends by frames that have already been approved, which converts an open ended generation into a constrained interpolation.
The value of this is mostly about review. An unbounded generated clip is difficult to approve because everything in it is still open: the subject, the light, the framing and the movement can all change. A shot bounded by two approved stills has only one variable left, which is how the transition behaves. Clients can review that quickly and studios can iterate on it cheaply, and the endpoints cannot drift away from the visual language that was agreed.
The technique is particularly strong for transformations, which are otherwise difficult to direct. A product changing state, a space changing time of day, a material forming, a logo assembling, a raw ingredient becoming a finished dish: in each case the two endpoints are the meaningful creative decisions and the movement between them is execution. Generating from both ends means the two decisions that matter are made as still images, where they are easy to judge.
It is also the practical way to build sequences that need to connect. If the last frame of one shot is used as the first frame of the next, the two shots join without a visible discontinuity in lighting, colour or position. Chaining shots this way produces continuous sequences considerably longer than a single generation can sustain, which is one answer to the shot length limits that constrain the format.
The failure modes are predictable and worth designing around. If the two endpoints are too dissimilar, the model has to invent a large amount of intermediate content and the result is a morph rather than a movement, which reads as unstable. If they are too similar, the shot has no motion worth watching. If an object present in the first frame is absent in the last, it usually dissolves in a way that draws attention to itself. The working rule is that the two frames should be plausibly connected by a single continuous action.
Duration interacts with endpoint distance. A short interval between dissimilar frames produces an abrupt, artificial transformation; a long interval between similar frames produces drift, where the model wanders because it has time to fill and little to do. Matching the generated duration to the amount of change required is a judgement that improves quickly with practice and is worth recording in the shot plan rather than rediscovering each time.
In practical use the technique fails in a particular way that catches people out: the last frame is often approximated rather than matched exactly. The generated final frame may be close to the supplied image without being identical, which matters when the next shot has been built to start from that exact image. The reliable workaround is to hold the intended frame briefly at the end of the shot, or to cut a frame or two earlier, rather than assuming a perfect landing.
Nour (2026), comparing prompt engineering with model selection, found the two to be distinct levers with different effects rather than interchangeable ones. First and last frame conditioning belongs to a third category, which is structural control, and in practice it outperforms both wording and model choice for any shot where the endpoints matter more than the style. Recognising which category a problem belongs to is what stops a studio from rewriting a prompt twenty times to fix something that conditioning solves in one attempt.
The workflow implication is that this technique rewards the stills first approach rather than working alongside it. If a project has already produced and approved a set of hero frames, those frames become endpoints, and much of the film can be assembled as transitions between things the client has already signed off. Projects that generate motion first and stills never have nothing to condition from and are back to describing shots in words.
For clients the useful thing to know is that asking to see the start and end frames of a shot is a reasonable and productive request. It is faster to review than a clip, it is the point at which changes are cheapest, and it is how a studio working this way is already thinking about the shot. A studio that cannot show you the endpoints is generating hopefully rather than deliberately.
References
Nour, R. R. (2026). Prompt engineering versus model selection for cognitive accessibility in large language models: An empirical study. IEEE Access, 14, 44740–44754. https://doi.org/10.1109/ACCESS.2026.3667133