Text to Video AI
Write what you want to see and get it back as footage. No camera, no stock library, no timeline — just a prompt, a model, and the aspect ratio you are going to publish in.
- Prompt to finished clip
- Choose the model per shot
- Reusable characters
- Vertical, square, and wide
How to turn text into video
The prompt does most of the work. These three steps are the whole loop, and the second one is where the quality difference usually comes from.
- 01
Write the shot, not the story
Describe one continuous shot: the subject, what it is doing, the camera move, the lighting, and the style. A prompt that describes a sequence of events gets a clip that tries to do all of them badly.
- 02
Choose a model and the output shape
Pick the model whose strengths match the shot — motion, faces, or stylisation — then set the length and aspect ratio. The form only offers combinations the model supports.
- 03
Generate, review, iterate
Runs happen in the background, so you can queue variations of the same prompt and compare them side by side in your library rather than waiting on each one in turn.
What makes a text-to-video run usable
Prompt guidance built in
The composer shows the model's own prompt length limit and flags the parts of a prompt models routinely ignore, so you find out before you spend credits rather than after.
Consistent characters
Reference a saved character in the prompt and the same face appears across every clip. This is the difference between a set of unrelated shots and a series that reads as one piece of content.
Model switching without rewrites
Your prompt, references, and settings carry across when you change model, so comparing two models on the same idea costs one click rather than a rebuild.
Queue several variations
Generation does not block the page. Submit three phrasings of the same shot, keep working, and judge them together when they land.
Finish in place
Upscale to 2K or 4K, reframe to another aspect ratio, or extend a clip that ends a beat too early — without exporting to another tool first.
Sound on the same page
Add generated voice-over, music, or effects from the audio studio, using the same credits and the same library.
Where text to video is the right input
Scenes that do not exist
A location you cannot shoot, a concept that has no footage, a stylised world — anything where there is nothing to photograph in the first place.
Hooks and openers
The three seconds at the top of a short-form video, generated in a dozen variants until one of them holds attention.
Storyboards and pitches
Moving boards that show a client the intent of a sequence long before anything is scheduled or booked.
B-roll and cutaways
Filler shots that match an existing edit's look, generated to the exact length the timeline needs.
Stylised explainers
Abstract or illustrative motion for concepts that would look flat as a slide or a stock clip.
Ad variants
The same message shot several ways, so a campaign can test more than one creative without a second production day.
Writing a prompt that survives the model
Most disappointing text-to-video results come from prompts that ask for a scene rather than a shot. A model generating four to twelve seconds cannot narrate a sequence of events; what it can do is hold one continuous moment convincingly. Prompts that name a single subject, one action, one camera behaviour, and a lighting condition are consistently better than prompts that describe a small film.
The second habit worth building is specificity about the frame rather than the story. "Slow dolly in on a ceramic mug on a windowsill, morning backlight, shallow depth of field" gives the model a composition to solve. "A cosy morning scene" gives it nothing to hold onto, and the result will drift differently on every run, which makes iteration feel random rather than directed.
Finally, treat consistency as a setup problem rather than a prompting problem. If the same person, product, or style has to appear in more than one clip, save it once — as a character or a reference image — and point at it. Re-describing a face in words gets you a different face every time, however careful the description is, because the words are not the anchor. The reference is.
Text to video FAQ
- Is text to video free?
- Yes, within your free credits. Every account starts with credits that work across video, image, and audio generation, and no card is required to spend them. Paid plans add more credits, higher resolutions, and watermark-free exports.
- How long does a generation take?
- Most runs finish in one to three minutes depending on the model, the length, and how busy the provider is. Generation runs in the background, so you can queue more prompts or close the tab and come back to the results.
- Can I keep the same character across several clips?
- Yes. Save a character from your own photos, then reference it in each generation. That is the reliable way to keep one face across a series — describing the same person in words produces a different person each run.
- What length and aspect ratio can I generate?
- Lengths generally run from four to twelve seconds per clip, and 9:16, 1:1, 4:5, and 16:9 are all available depending on the model. The available combinations are shown in the form before you submit.
- Can I use the results commercially?
- Yes, under the terms of service, on a paid plan. You remain responsible for what you put into a prompt — in particular for not generating a real person's likeness without their agreement.
Other ways into the same studio
Try a prompt and see what comes back
Free credits on sign-up, every model in the catalog, and results that land in a library you can keep working from.