Field guide
Text-to-Video vs Image-to-Video: Which Should You Use?
Choose text-to-video or image-to-video from your brief, compare two ways to build the same shot, and decide when to change your source or switch modes.
The practical difference
Text-to-video starts with a written scene. Image-to-video starts with an image plus directions for how the scene should develop. If you need a new setting for a campaign teaser, start from text. If you have already chosen the portrait, outfit, or illustration, start from that image.
An image provides the visual starting point; text leaves more of that appearance to be generated. These input responsibilities are described in Runway's text-to-video and image-to-video documentation. Both guides concern Gen-4.5.
Choose a starting mode
Scroll sideways to compare all columns →
| Your requirement | Start with | What to resolve first |
|---|---|---|
| Explore an imagined setting with no approved artwork | Text-to-video | Describe the setting and the action you want to see |
| Animate a chosen creator portrait or illustration | Image-to-video | Choose a source you are allowed to use, in a useful composition |
| Keep a specific product's visible design | Image-to-video | Check the source details, then inspect whether the output preserves them |
| Create atmospheric footage where exact appearance is flexible | Text-to-video | Define the event, framing, and visual style |
| Show a large movement beyond a tightly cropped photograph | Reconsider the source or try text-to-video | Decide whether matching the photograph or revealing a new scene matters more |
| Preserve a character across several scenes | Plan references and continuity first | Neither input mode alone guarantees identity across clips |
| Animate only a still image's position or scale exactly | Consider conventional editing | You may need precise transforms rather than newly generated motion |
One brief, two ways to build it
Consider this brief: a red paper boat drifts across a rain puddle while the camera stays low beside the water. These original drafts show what you would prepare for each route.
The text-to-video route
You can begin with the shot itself:
A low view of a red paper boat in a shallow puddle on gray pavement. The boat drifts slowly from left to right, making small ripples. The camera remains beside the puddle. Soft overcast daylight.
This route asks the model to invent the boat, puddle, framing, and motion together. Decide before generating which details can vary. If the boat's exact fold or the pavement pattern is unimportant, several different starting scenes may satisfy the brief.
Review appearance and action separately: did it create a plausible paper boat, and did that boat actually drift? Use the AI video prompt guide to revise the particular mismatch.
The image-to-video route
Start with your chosen boat image, then describe the change:
The boat drifts slowly from left to right across the puddle, making small ripples. The camera stays in place beside the water.
Now the choice of image is part of the creative work. Does it show enough space for the boat to move right? Is the boat already cut off? Is the view low enough for the brief? If those answers are wrong, the motion prompt is being asked to solve a composition problem too.
Follow the image animation walkthrough for the complete workflow, and the image preparation guide for source selection.
What our archived API examples show
Two retained site examples show the input difference. The Stallion text-to-video request supplied text with a duration, resolution, aspect ratio, and audio setting. The Paper Garden image-to-video request supplied text, image references, duration, resolution, and audio settings.
See the horse outputs and the Paper Garden input and result to follow each request through to its result.
Example details
These Seedance 2.5 provider API runs were made on August 27, 2026. Paper Garden used the same image as both first and last reference; it is not a first-image-only example. The subjects differ, so the pair illustrates input responsibilities rather than comparing mode quality. For a new generation, use the controls available in the current tool. See our editorial method.
When to switch modes
Switch from text-to-video to image-to-video when you like the motion idea but need a particular starting appearance. For example, if a campaign already has an approved illustration, use that source rather than repeatedly describing its colors and layout from scratch. Then check the video for changes to the details that matter.
Switch from image-to-video to a different image or to text-to-video when the source conflicts with the shot. A close-up of the boat's bow is a poor starting brief for a wide view of the entire puddle. You can choose a wider source, reduce the requested reveal, or allow the scene to be generated from text. For a hero background, a newly imagined setting may work; for an existing portrait, keeping the chosen appearance may matter more than the reveal.
If the main problem is a face or outfit changing during a clip, consult the character consistency guide. Compare those details at the beginning, middle, and end of the new result.
Compare the total work and cost
Text-to-video lets you begin without preparing an image. Image-to-video makes use of visual work you already have. To compare the cost of a finished shot, include preparation and every generation you needed.
For a fair personal comparison, record image preparation time, generation settings, the cost of each attempt, and how many results meet your brief. Count rejected clips too, using the same requirements for both routes.
Use the price shown for each selected configuration and the Credits explanation. Compare the displayed cost of each complete configuration.
Make the first choice from your brief
Write down the one thing the result must preserve. If it is a particular existing frame, start with Image to Video. If it is an event you want to invent, start with Text to Video. Save the first result even when it misses the brief: it gives you a concrete basis for changing the prompt, the source, or the mode.
Use the method