You upload a beautiful still image, type "make it move", and get back four seconds of a slow zoom with nothing happening. Almost everyone hits this wall on their first day with image-to-video AI.
The problem is rarely the model. It is the prompt. When you start from an image, the model already knows what the scene looks like. What it does not know is what should change. This guide shows you how to describe that change so your stills turn into clips with real motion.
Why Image-to-Video Prompts Are Different
A text-to-video prompt has to describe everything: the subject, the setting, the light, the style, and the motion. An image-to-video prompt has a much smaller job, because the first frame already answers most of those questions.
That sounds easier, but it causes a common mistake. People describe the image again, word for word. The model reads "a woman in a red coat standing on a bridge at sunset" and concludes that the best result is the picture it was given, with as little change as possible.
The rule is simple: do not describe what the model can see. Describe what it cannot see yet. That means motion, timing, camera behavior, and sound.
The Four Things Your Prompt Must Answer
Every strong image-to-video prompt answers four questions.
1. What moves?
Name the subject that moves and the action it takes. Use a concrete verb. "The woman turns her head toward the camera" works. "The woman is dynamic" does not.
2. How does it move?
Speed and quality of motion matter as much as the action. Slowly, suddenly, gently, in one smooth motion, with a small stumble. These words separate a lifelike clip from a robotic one.
3. What does the camera do?
If you say nothing, most models default to a gentle push-in. Decide on purpose: a static locked-off shot, a slow dolly forward, a pan to the left, an orbit around the subject, or a handheld feel.
4. What else changes in the world?
Secondary motion sells realism. Hair moves in the wind, steam rises from a cup, leaves drift past, city lights flicker on. One or two of these details make the frame feel alive.
A Reusable Template
Here is a template you can copy for most shots:
Subject + action verb + manner of motion. The camera + camera move. One or two secondary motions. Optional sound.
For example:
The old fisherman lifts his fork and takes a slow bite, smiling as he chews. The camera holds a static medium shot. Gulls glide across the harbor behind him and the water ripples softly. Quiet harbor ambience with distant gull calls.
Notice what is missing: no description of his hat, his sweater, or the pasta. The image already contains all of that.
12 Image-to-Video Prompt Examples
Use these as starting points. Swap the subject for whatever is in your image.
Portraits and people
- Natural reaction: "She glances down at her phone, then looks up at the camera and breaks into a wide smile. Static shot. A light breeze moves a few strands of her hair."
- Walking shot: "He starts walking toward the camera at a relaxed pace. The camera slowly dollies backward to keep him centered. Pedestrians pass in the soft-focus background."
- Talking presenter: "The presenter speaks to the camera with small, natural hand gestures and occasional nods. Static medium shot. Subtle head movement, natural blinking."
Products
- Hero rotation: "The bottle rotates slowly clockwise on the turntable. The camera stays locked. Soft light glides across the glass and catches the label."
- Pour shot: "Coffee pours into the cup in one smooth stream, and steam begins to rise. Slow push-in toward the rim of the cup."
- Unboxing reveal: "The lid of the box lifts open, revealing the watch inside as warm light spills out. The camera tilts down slightly."
Landscapes and places
- Time passing: "Clouds race across the sky in a time-lapse while the shadows on the mountain shift. The camera holds a wide static frame."
- Drone reveal: "The camera rises slowly above the treetops to reveal the lake behind them. Morning mist drifts across the water."
- City at night: "Car headlights stream along the avenue below, and windows in the towers flicker on one by one. Slow pan to the right."
Stylized and abstract
- Illustration to life: "The painted fox blinks, flicks its tail, and turns its head toward the moon. Snowflakes drift through the frame. Keep the watercolor style."
- Product grid: "The rows of colorful blocks rise and fall in a wave that travels from front to back. Low camera angle, slow forward glide."
- Logo moment: "The logo emerges from a burst of light particles that settle into place. Static camera, gentle glow pulse at the end."
Common Mistakes and How to Fix Them
| Problem | Likely cause | Fix |
|---|---|---|
| Almost no motion | The prompt describes the image instead of the action | Remove scene description and add one clear action verb |
| The face changes or warps | Too much motion for a close-up | Ask for smaller motions and a static camera |
| Chaotic camera | Several camera moves in one prompt | Pick one camera move per clip |
| The style drifts | The model "improves" the look | Add "keep the original style and colors" |
| Motion looks robotic | No manner of motion | Add adverbs such as slowly, smoothly, gently |
Choosing the Right Starting Image
A better input image saves you more time than a better prompt.
- Leave room for motion. If your subject will walk forward, give them space in the frame. If the camera will pan right, leave space on the right.
- Avoid hard crops at the edges. A hand cut off by the frame edge is hard for the model to animate cleanly.
- Match the aspect ratio to the output. A vertical image for a vertical clip avoids awkward cropping.
- Use clear, sharp images. Blur and heavy compression become flicker and smearing in video.
Using First and Last Frames
Some models, such as Veo 3.1 on veo4.dev, accept both a first frame and a last frame. You control where the clip starts and where it ends, and the model fills in the motion between them.
Use it for transformations: a closed flower and an open flower, an empty room and a furnished room, a product in its box and the same product on the table. Keep the two frames consistent in lighting and camera angle. Large jumps between them produce morphing instead of motion.
Try It Yourself
The fastest way to learn is to take one image and run three prompts: one with only an action, one with an action and a camera move, and one with everything from the template. Compare the results side by side and you will see the difference immediately.
You can upload a reference image and test these prompts on the Veo 4 AI video generator. For more prompt ideas, read our Veo 4 prompts guide and the 100 prompt examples.