Skip to content

Image to Video AI

Upload a photo and get motion back. The image can set the opening frame or just define the subject — either way the result keeps the thing you actually photographed.

  • First frame or subject reference
  • Keeps faces and products recognisable
  • Any aspect ratio you publish in
  • Works from a phone photo

How to turn a photo into video

Starting from an image removes the hardest part of prompting: the model no longer has to guess what the subject looks like.

  1. 01

    Upload the photo

    Any clear image works — a product shot, a portrait, a landscape, a screenshot of a design. Higher resolution helps, but a phone photo is usually enough.

  2. 02

    Say what should move

    Add a short prompt describing the motion you want: a camera push, a turn of the head, wind through a field. The subject stays; the prompt only has to explain the movement.

  3. 03

    Choose the model and generate

    Models differ in how much they preserve the source frame. Pick one, set the length and aspect ratio, and the run happens in the background while you keep working.

Why starting from an image works better

The subject is already decided

A prompt has to describe a face or a product in words and hope. An image hands the model the exact thing, which is why image-to-video usually needs fewer runs to land.

First frame or reference

Use the photo as the literal opening frame when the composition matters, or as a subject reference when you want the model free to choose the framing.

Product-safe motion

Commerce and marketing runs keep labels, shapes, and colours intact while the camera moves around them — the point of the clip is the product, not a reinterpretation of it.

Multi-image input

Some models accept several references at once, so a subject, a background, and a style can each come from a different photo.

Characters from your own photos

Save a set of photos as a character and every later generation can call on it, instead of re-uploading the same face for each clip.

Finish without leaving

Upscale, relight, reframe, or extend the result in the same workspace, on the same credits, in the same library.

What people animate

Product photos

A single catalog image becomes a short rotating or pushing shot that can run on a listing or as an ad.

Portraits

A still portrait gets a subtle head turn, a blink, or a camera move, without the face becoming somebody else.

Old photographs

Archive and family photos get gentle motion and depth while keeping their original framing and grain.

Design mockups

A flat mockup or screenshot becomes a moving presentation shot for a pitch deck or a launch post.

Illustrations and art

Drawn or generated artwork gets parallax and motion while the linework and palette stay intact.

Real estate and travel

A property or location photo turns into a slow move that reads far better in a feed than a static image.

Choosing between first frame and reference

Image-to-video models accept a photo in two quite different ways, and picking the wrong one is the most common reason a run disappoints. Passing the image as the first frame means the clip literally begins on that composition and moves away from it. Everything you framed — crop, background, lighting — is preserved, and the model's only job is motion. That is what you want for a product, a logo, or any shot where the composition is the deliverable.

Passing the image as a subject reference is a looser contract. The model keeps who or what is in the picture but is free to reframe, relight, and place the subject in a new context. That is the right choice when the source photo is evidence of a subject rather than a composition you care about — a face you want in a different scene, a product you want on a different surface.

In practice the fastest workflow is to decide which of the two you need before you write the prompt, because the prompt should say different things in each case. As a first frame, the prompt describes only movement. As a reference, the prompt has to describe the new scene as well, and the image is only anchoring the subject inside it.

Image to video FAQ

What kind of photo works best?
A sharp, well-lit image where the subject is clearly separated from the background. Heavy blur, very low resolution, or a subject cut off at the edge of frame all make the motion less predictable. A normal phone photo is fine.
Will the face still look like the person in the photo?
Yes, that is the point of the mode, though how strictly the likeness is preserved varies between models. Saving the photos as a character gives the most consistent result across several clips.
Can I upload more than one image?
Some models accept multiple reference images — for example a subject and a background. The form shows how many images the selected model takes, and whether any of them are required.
How long is the generated clip?
Typically four to twelve seconds depending on the model. If you need longer, extend the clip or generate a second shot and join them.
Whose photos am I allowed to upload?
Your own, or photos of people who have agreed to appear. Publishing a generated likeness of someone who has not agreed is not permitted, and anyone can ask for a published post of themselves to be taken down.

Related tools and inputs

Upload one photo and see it move

Free credits on sign-up, no card required, and the result lands in a library you can upscale, reframe, or extend from.