✦ New prompt drops are live — browse the latest Veo experiments
V/

Guide / 6 min read

Veo 3 Prompt Structure: The 6 Elements That Actually Matter

Most Veo 3 prompts fail for the same reasons. This breaks down the six-part structure that produces consistent results, with before-and-after examples.

A prompt that names a subject gets you a generic result. A prompt that describes a shot gets you something usable. The difference is structure, and it comes down to six elements.

1. Subject

Concrete visual detail, not category. "A man" is a category. "A man in his sixties, weathered face, wearing a faded denim jacket" is a subject the model can render consistently.

2. Action

One clear action that fits in the clip duration. You have roughly eight seconds. "Walks across the room, sits down, opens a laptop and starts typing" is four actions crammed into a window that fits one. Pick the one that carries the shot.

3. Camera

This is the element most people skip, and it is the one that most separates amateur output from professional. Specify shot size and movement:

  • Shot size — extreme close-up, close-up, medium, wide, extreme wide.
  • Movement — locked-off, slow push in, dolly out, handheld tracking, crane up, orbit.
  • Lens feel — "35mm", "shallow depth of field", "wide angle with slight distortion".

"Cinematic" is not a camera instruction. "Slow dolly in, medium close-up, shallow depth of field" is.

4. Setting

Location and time of day. Time of day does double duty because it implies lighting direction and colour temperature.

5. Lighting

Name the source and quality. "Golden hour backlight with lens flare", "harsh overhead fluorescent", "single practical lamp, deep shadows", "soft overcast daylight". Lighting is most of what makes a shot read as expensive or cheap.

6. Audio

Veo 3 generates sound. If you do not specify it, the model guesses. Separate the layers:

  • Ambient — room tone, weather, traffic, crowd.
  • Specific effects — footsteps on gravel, a door latch, paper tearing.
  • Dialogue — in quotes, attributed to a speaker.

Putting it together

Before:

A chef cooking in a kitchen, cinematic.

After:

Medium close-up, slow dolly in. A chef in his forties, sleeves rolled, sears a steak in a cast iron pan — flames briefly rise. Professional kitchen, late evening. Hard key light from overhead, deep shadows, steam catching the light. Audio: aggressive sizzle, extractor fan hum, distant kitchen chatter.

Order matters less than completeness

The model does not require this exact sequence. What it requires is that the information is present. A prompt missing camera and lighting will produce an inconsistent result no matter how the words are arranged.

One change at a time

When a generation is close but wrong, change a single element and regenerate. Changing three things at once means you learn nothing about which one mattered. This is slower for one clip and dramatically faster across a project.

See prompts that worked

Every prompt in the library comes with the video it produced, so you can judge the result before you spend a credit.

Browse the library

Keep reading