Skip to content

Frame to Video AI

Two images, one clip. Set the opening frame and the closing frame and the model generates the movement that connects them — so the composition you end on is the one you chose, not the one the model drifted into.

  • Start and end frames you control
  • Predictable, repeatable transitions
  • Loop a clip by reusing one frame
  • Works with generated or shot stills

How first and last frame generation works

This mode swaps prompting for art direction. You are not asking for a scene — you are handing over the two frames that bracket it and letting the model solve the middle.

  1. 01

    Choose the frame you open on

    Upload or generate the still that starts the clip. Composition matters more than resolution here: whatever is in this frame is what the first moment of the video looks like, and the model will not reinterpret it.

  2. 02

    Choose the frame you land on

    Add the still the clip should arrive at. The bigger the difference between the two frames, the more invention the model has to do — and the more likely the middle wanders. Small, deliberate changes read best.

  3. 03

    Describe the movement, not the subject

    The prompt in this mode is about how to get from A to B: a camera push, a turn of the head, a product rotating, light changing. The subject is already established by the frames.

Why bracket a shot instead of prompting it

The ending is decided in advance

Text-to-video gives you whatever composition the model drifts into by the last frame. Bracketing means the final frame is a decision, which matters when the clip has to cut into something else.

Clean loops

Use the same image as the first and last frame and the clip returns to where it began, which is the straightforward way to build a loop that does not visibly jump.

Transitions between two shots

Take the last frame of one clip and the first frame of the next, and generate the movement that joins them — a transition that belongs to the footage rather than an effect laid on top.

Product turns and reveals

Front view to three-quarter view, closed to open, empty to full: two frames pin both ends of the motion so the product stays itself throughout.

Stills you already have

Photographs, renders, screenshots, or images generated in the image studio all work as frames. Nothing has to be produced inside one tool.

Reruns that stay comparable

Because both ends are fixed, running the same pair again with a different prompt changes only the motion — which makes iterating on movement an actual experiment rather than a reroll.

Where bracketing earns its keep

Looping backgrounds

Ambient motion for a site header or a stream overlay that has to run indefinitely without a visible seam.

Before and after

A room, a face, a garment, a chart — two states that mean something as a pair, animated into a single continuous change.

Logo and title moves

A brand mark that arrives into its lockup, ending on the exact frame the rest of the edit expects.

Match cuts

Two compositions that rhyme, joined by generated motion instead of a cross-dissolve.

Product rotations

A controlled turn between two catalog angles, so the shot ends on the hero image already used elsewhere.

Storyboard to animatic

Boards drawn as key frames, generated into the movement between them to show timing to a client.

Choosing two frames that the model can actually connect

The failure mode of frame-to-video is asking for too much change. If the opening frame is a wide exterior at night and the closing frame is a close-up of a face indoors, there is no continuous motion that connects them — so the model invents a cut, a morph, or a smear, and none of those are what you wanted. Treat the two frames as the ends of one camera move or one action, and the results stay coherent.

Consistency of everything except the intended change is what makes the middle believable. Keep the lighting, the lens, the colour treatment, and the background the same in both frames, and vary only the thing that is supposed to move. When the two stills come from different sources — one photographed, one generated — that consistency is worth spending a few minutes on before generating, because the model will otherwise spend the whole clip trying to reconcile two looks.

Length is the other lever. A four-second clip between two very similar frames produces slow, controlled motion; the same pair over twelve seconds gives the model room to add movement you did not ask for. If a result feels padded, shortening the clip is usually a better fix than rewriting the prompt.

Finally, remember that the frames are an anchor, not a straitjacket for style. The prompt still controls how the movement feels — a handheld drift and a locked-off dolly are different requests even with identical frames — so it is worth running the same pair two or three ways before settling.

First and last frame FAQ

Do I need both frames?
No. With only a first frame this behaves like ordinary image-to-video, and the ending is left to the model. The second frame is what turns it into a controlled move, so add it whenever the final composition matters.
Can the two frames be different sizes?
They should share an aspect ratio. Mismatched shapes get cropped or letterboxed to the output ratio you pick, which usually means part of one frame is quietly discarded.
How do I make a seamless loop?
Use the same image for both frames. The clip leaves and returns to that composition, so it can be played end to end without a visible cut.
Which models support it?
The mode is offered only for models that accept a start and an end image, and the form shows just those. If a model you want is missing, it does not support bracketing yet.
Is it more expensive than text to video?
No. Cost is driven by the model, the length, and the resolution — the same as any other run. The frames themselves cost nothing.

Other ways into the same studio

Pick two frames and see what connects them

Free credits on sign-up, every compatible model in the catalog, and results that land in a library you can keep working from.