AI video prompting: Veo 3.1, Sora 2 migration and a reusable shot formula
Write clearer AI video prompts with a practical shot formula, six editable examples, Veo 3.1 guidance and current Sora 2 migration dates.
Availability update before you choose a model
This guide keeps Sora 2 in the title because people still need its prompt and migration information, but it is no longer a safe choice for a new long-term workflow. OpenAI says the Sora web and app experiences ended on April 26, 2026 and the API is scheduled to end on September 24, 2026. OpenAI's model catalogue labels Sora 2 as legacy.
Google's current Gemini API video overview recommends Gemini Omni Flash as the default for general video generation and positions Veo 3.1 for capabilities such as scene extension, last-frame control and legacy-pipeline integration. Check the live provider documentation, region, account and pricing before designing a production workflow.
The answer-first video prompt formula
A useful starting structure is: cinematography + subject + action + context + style and ambience. This follows Google's published Veo 3.1 prompting framework. Add a sixth block for audio when the model supports it, and keep API controls such as duration, size and reference files in their proper settings rather than hiding them in prose.
- Cinematography: framing, viewpoint and one primary camera behaviour.
- Subject: the person, object or place, described with stable facts.
- Action: what changes during the clip, in a clear sequence.
- Context: setting, time, weather, foreground and background.
- Style and ambience: lighting, palette, texture, pace and emotional tone.
- Audio, when supported: dialogue, sound effects and ambient sound.
Write motion as a timeline, not a pile of adjectives
Video prompts need change over time. Name the starting state, the subject's action, the camera response and the intended end state. Simple instructions are easier to diagnose than several simultaneous moves. Complex movement is not forbidden, but it should be staged deliberately and tested rather than presented as universally reliable.
- Locked shot: useful when subject action or dialogue is already complex.
- Dolly or tracking move: state direction, speed and what remains framed.
- Pan or tilt: name the reveal and where the move stops.
- Orbit: specify the portion of the arc and preserve a clear subject anchor.
- Multi-shot sequence: use provider-supported timestamp or storyboard controls when available.
Duration and controls are model-specific
As of this review, Google's Veo 3.1 guide documents 4, 6 and 8-second clips, while OpenAI's video API reference lists 4, 8 and 12 seconds for Sora 2 jobs. Those are request options, not a rule that every creative prompt should contain. Model versions and controls change, so avoid evergreen claims such as "8–12 seconds is always best."
Character and scene consistency without false guarantees
Repeat stable facts—approved appearance description, wardrobe, key props, environment and lighting direction—across related shots. Where supported, use the provider's reference-image, first/last-frame, ingredient, edit, extension or remix controls. Do not assume that two differently generated clips will preserve identity exactly.
Review every result for face drift, altered logos, changed text, extra objects and continuity errors. For a person's likeness, obtain consent and avoid deceptive or harmful use. Keep a record of source assets and model settings so a result can be reproduced or corrected.
Audio instructions for Veo 3.1
Google documents native audio in Veo 3.1. Put exact dialogue in quotation marks, describe sound effects clearly and separate ambient noise from music. If silence matters, say so. Generated speech and lip movement still need review for words, timing, accents and unintended background sound.
Six editable prompt examples
These are teaching examples, not "battle-tested" performance claims. Replace the subject, setting and controls, then compare variations in the model you use.
Slow portrait reveal
Medium portrait shot of an adult ceramic artist beside a kiln in a quiet workshop. The artist brushes dust from a finished bowl and looks toward the window. Slow, steady dolly-in; soft morning light from camera-left; warm clay colours; restrained documentary mood. Ambient sound: light room tone and the soft scrape of pottery tools.
Veo 3.1 product clip
Macro close-up, a condensation-covered glass bottle of cold brew on dark walnut. One droplet travels down the label while the camera performs a slow lateral slide. Soft window light from the right, deep controlled shadow on the left, commercial realism, muted earth tones. SFX: one quiet drop and subtle café ambience.
Landscape reveal
Wide aerial tracking shot moving left to right above a misty pine forest at sunrise. Foreground trees create parallax while distant mountains gradually appear. Smooth motion, wide-angle perspective, warm sunlight through fog, natural colour, calm travel-documentary atmosphere.
Presenter with dialogue
Medium close-up of an adult presenter in a navy jacket speaking to camera in a modern office. Locked camera at eye level, soft key light from camera-left, gentle background blur. The presenter says, "Here is the result of today's test." Clear speech, low office room tone, no background music.
Controlled action shot
Low-angle tracking shot following an adult runner crossing a rooftop training course at dusk. The runner clears one short gap and lands in a stable crouch. Wide-angle perspective, stabilized motion, orange sky with cool shadows, realistic momentum, no additional camera move.
Food detail loop concept
Top-down shot of a steaming bowl of ramen on dark slate. The camera makes one slow partial orbit while steam rises and a cook places chopsticks beside the bowl. Soft overhead light, natural food texture, restrained restaurant-commercial styling. End with framing close to the opening composition for easier editing into a loop.
A practical evaluation checklist
- Did the main subject remain recognizable and physically coherent?
- Did the intended action happen in the right order?
- Did framing and camera direction match the request?
- Did text, logos, hands and small objects mutate?
- Did dialogue and sound match the script without unwanted additions?
- Would a shorter prompt or a separate shot make the failure easier to fix?
Change one variable per iteration and save the prompt, model version, request settings and reference assets. That small habit produces more useful evidence than adding more adjectives.
Frequently asked questions
Is Sora 2 still available in 2026?
OpenAI discontinued the Sora web and app experiences on April 26, 2026 and says the Sora API will be discontinued on September 24, 2026. Existing API users should confirm access and plan a migration.
What is a reliable AI video prompt structure?
Start with cinematography, subject, action, context and style or ambience. Add dialogue, sound effects or ambient audio only when the selected model supports audio generation.
How long should an AI video prompt be?
There is no universal word limit. Use enough detail to remove ambiguity, then delete conflicts and low-priority adjectives. Provider controls such as clip duration belong in the request settings when the API exposes them.
How do I improve character consistency?
Reuse the same approved character description and reference assets, keep wardrobe and scene facts stable, and change one shot variable at a time. Consistency is model-dependent and should not be promised as exact.
Primary sources and review date
Checked July 28, 2026 against the OpenAI Sora discontinuation notice, OpenAI video API reference, OpenAI Sora 2 model page, Google Gemini API video overview and Google Cloud's Veo 3.1 prompting guide.
Build an editable video prompt
Choose a model, define the shot, then verify current provider controls before generation.
Open Video Prompt Generator →