Guide / 6 min read
Veo 3 Prompt Structure: The 6 Elements That Actually Matter
Most Veo 3 prompts fail for the same reasons. This breaks down the six-part structure that produces consistent results, with before-and-after examples.
A prompt that names a subject gets you a generic result. A prompt that describes a shot gets you something usable. The difference is structure, and it comes down to six elements.
1. Subject
Concrete visual detail, not category. "A man" is a category. "A man in his sixties, weathered face, wearing a faded denim jacket" is a subject the model can render consistently.
2. Action
One clear action that fits in the clip duration. You have roughly eight seconds. "Walks across the room, sits down, opens a laptop and starts typing" is four actions crammed into a window that fits one. Pick the one that carries the shot.
3. Camera
This is the element most people skip, and it is the one that most separates amateur output from professional. Specify shot size and movement:
- Shot size — extreme close-up, close-up, medium, wide, extreme wide.
- Movement — locked-off, slow push in, dolly out, handheld tracking, crane up, orbit.
- Lens feel — "35mm", "shallow depth of field", "wide angle with slight distortion".
"Cinematic" is not a camera instruction. "Slow dolly in, medium close-up, shallow depth of field" is.
4. Setting
Location and time of day. Time of day does double duty because it implies lighting direction and colour temperature.
5. Lighting
Name the source and quality. "Golden hour backlight with lens flare", "harsh overhead fluorescent", "single practical lamp, deep shadows", "soft overcast daylight". Lighting is most of what makes a shot read as expensive or cheap.
6. Audio
Veo 3 generates sound. If you do not specify it, the model guesses. Separate the layers:
- Ambient — room tone, weather, traffic, crowd.
- Specific effects — footsteps on gravel, a door latch, paper tearing.
- Dialogue — in quotes, attributed to a speaker.
Putting it together
Before:
A chef cooking in a kitchen, cinematic.
After:
Medium close-up, slow dolly in. A chef in his forties, sleeves rolled, sears a steak in a cast iron pan — flames briefly rise. Professional kitchen, late evening. Hard key light from overhead, deep shadows, steam catching the light. Audio: aggressive sizzle, extractor fan hum, distant kitchen chatter.
Order matters less than completeness
The model does not require this exact sequence. What it requires is that the information is present. A prompt missing camera and lighting will produce an inconsistent result no matter how the words are arranged.
One change at a time
When a generation is close but wrong, change a single element and regenerate. Changing three things at once means you learn nothing about which one mattered. This is slower for one clip and dramatically faster across a project.
See prompts that worked
Every prompt in the library comes with the video it produced, so you can judge the result before you spend a credit.
Browse the libraryKeep reading
How to Generate Videos With Veo 3 (Step-by-Step Guide)
A practical walkthrough of generating AI video with Google Veo 3 — where to access it, how to structure a prompt, what the model does well, and the mistakes that waste credits.
How to Use Google Flow for AI Video (Complete Walkthrough)
Google Flow is the filmmaking interface built on Veo 3. Here's how the tools work — Frames to Video, Ingredients, Scene Builder and Extend — and how to build a multi-shot sequence that stays consistent.
Veo 3 vs Sora vs Kling: Which AI Video Model to Use
An honest comparison of the three major AI video models — where each one wins, where each falls down, and how to pick based on what you're actually making.