Sora 2 & Veo 3 Prompting Guide: How to Write AI Video Prompts That Actually Work
How to write effective AI video prompts for Sora 2 and Veo 3 — camera moves, temporal consistency, shot grammar, and 30+ ready-to-copy templates.
Why Sora 2 and Veo 3 need different prompts than image models
Image models like Midjourney and Nano Banana describe a single frozen moment. Video models like OpenAI Sora 2 and Google Veo 3 have to plan motion, physics, lip sync, and temporal consistency across 5–60 seconds. That changes how you write the prompt: you are no longer describing a picture — you are directing a shot.
The 7-block prompt formula that works for both Sora 2 and Veo 3
- SHOT — frame size and angle (wide, medium, close-up, low-angle, top-down).
- SUBJECT — who or what, described with concrete physical detail.
- ACTION — one verb. Avoid chaining actions in a single clip.
- CAMERA — lens, movement, speed. One move per clip.
- LIGHTING — key direction, quality (soft/hard), color temperature.
- ENVIRONMENT — location, time of day, weather, background depth.
- STYLE — film stock, grade, era, reference (e.g. Kodak 500T, A24, anime cel).
Camera movement: what Sora 2 and Veo 3 can actually do
Both models render simple, named camera moves reliably. Stick to one move per clip and use the standard film vocabulary — the models were trained on it.
- Dolly in / out — physically moves toward or away from subject. Best for emotional emphasis.
- Track / parallax — moves sideways past subject. Great for reveals and landscapes.
- Orbit — circles the subject. Best at slow speed; fast orbits warp geometry.
- Crane / boom — vertical lift. Use sparingly; combine with locked horizon.
- Locked tripod — no movement. The most stable, highest-fidelity option. Underrated.
Avoid: "dolly in then crane up then orbit" in a single prompt — Sora 2 and Veo 3 will average the moves and you'll get drift and morphing geometry. One move per shot.
Temporal consistency: keeping the same face, outfit, and place across shots
The single biggest weakness of current AI video is identity drift between clips. Use these rules to lock it down:
- Repeat the subject description verbatim in every clip prompt — same age, hair, eye color, wardrobe, ethnicity, accessories.
- Anchor the lens — keep the same focal length (e.g. 35mm) across the sequence.
- Anchor the lighting direction — if key is camera-left in shot 1, keep it camera-left in shot 2.
- Anchor the color grade — name the same film stock or LUT in every prompt ("Kodak 2383", "teal-orange", "bleach bypass").
- Use Sora 2's "remix" or Veo 3's "extend" instead of regenerating from scratch when you need continuity.
Sora 2 specifics
- Prefers single-paragraph natural language over bullet structure.
- Sora 2 reads audio cues — name dialogue, sound effects, and music mood inside the prompt.
- Best clip length: 8–12 seconds. Longer clips lose subject coherence.
- Strong at physics (water, cloth, hair) when you name the material explicitly.
Veo 3 specifics
- Prefers block-structured prompts (SHOT / SUBJECT / ACTION / CAMERA…).
- Veo 3 has native synchronized audio — always include an AUDIO line.
- Strongest at talking heads and lip sync — name the dialogue in quotes.
- Best at locked tripod and slow dolly; less reliable on fast handheld.
30+ ready-to-copy prompt templates
Below are battle-tested templates. Swap the subject and grade to fit your project.
Cinematic dolly-in portrait (Sora 2)
Cinematic medium shot of a young woman standing on a rain-soaked Tokyo street at night, neon reflections on wet asphalt. Slow dolly-in, 35mm lens, shallow depth of field. Soft key light from the left, magenta and cyan rim light. Film grain, Kodak 500T look, 24fps.
Veo 3 structured product shot
SHOT: Macro top-down SUBJECT: A glass bottle of cold brew coffee with condensation ACTION: A single drop slides down the bottle CAMERA: Locked tripod, 60mm macro, f/2.8 LIGHTING: Soft window light from the right, deep shadow on the left ENVIRONMENT: Dark walnut table, scattered coffee beans STYLE: Commercial, photoreal, muted earth tones AUDIO: Subtle ambient room tone, soft drip
Parallax landscape reveal
Wide aerial parallax shot tracking left to right over a misty pine forest at sunrise, distant mountains revealed behind the trees. Smooth drone motion, 24mm lens, golden warm grade, volumetric god rays through fog.
Talking head, broadcast style
Medium close-up of a 40-year-old male presenter in a navy suit speaking directly to camera in a modern glass-walled office. Locked tripod, 50mm lens, eye-level. Soft key from camera-left, subtle fill, slight background bokeh. Natural lip sync, calm confident tone.
Action chase (low angle)
Low-angle tracking shot following a parkour runner leaping between rooftops at dusk. Handheld stabilized camera, 18mm wide lens, fast shutter, dynamic motion blur on edges. Warm orange sky, cool blue shadows, gritty cinematic grade.
Food hero loop
Top-down 360° orbit around a steaming bowl of ramen on a dark slate plate. Slow rotation, 50mm, soft overhead window light, visible steam rising, chopsticks resting on the rim. Photoreal, restaurant commercial look. Seamless loop.
Common mistakes that break Sora 2 and Veo 3 outputs
- Too many subjects. More than 2 named characters → identity collapse.
- Chained camera moves. One move per clip, always.
- Vague style words. "Cinematic" alone means nothing. Name the film stock or director.
- Conflicting lighting. "Golden hour at midnight" — pick one.
- Over-long prompts. Past ~120 words, Sora 2 starts ignoring tail tokens.
FAQ
How do I write a prompt for Sora 2?
Sora 2 prompts work best as a single paragraph that names the subject, the action, the camera move, the lens, the lighting, and the style. Put the most important elements first and keep total length under 120 words.
What is the best prompt structure for Veo 3?
Veo 3 responds to structured prompts: SHOT, SUBJECT, ACTION, CAMERA, LIGHTING, ENVIRONMENT, STYLE, AUDIO. Keeping each block on its own line gives Veo 3 the strongest temporal consistency.
How do I keep a character consistent across Sora 2 or Veo 3 shots?
Lock identity by repeating the exact same facial and wardrobe description in every shot, anchor lighting direction, and use the same lens (e.g. 35mm) and color grade across all clips.
What camera moves work best in AI video?
Slow dolly-in, parallax track, low-angle push, orbit, and locked tripod are the most reliable. Avoid chained moves (e.g. 'dolly then crane then orbit') in a single clip — they break temporal consistency.
Generate Sora 2 & Veo 3 prompts free
Unlimited variations. Camera-move library built in. No signup.