Published 2026-09-16 · Updated 2026-09-16

Gemini Omni 1.1 Flash: 40-Second Scenes, First/Last Frame and 4K

Short answer: Omni 1.1 Flash is Google's control update. Write the shot as a plain, timed description, mark your images with the tags the model reads (<FIRST_FRAME>, <LAST_FRAME>, <IMAGE_REF_1>), put events and lines on timecodes such as [0-3s], draft at 360p, extend in 10-second steps up to 40 seconds, then upscale the approved take to 1080p or 4K. There are no negative prompts and no system instructions - everything you want or forbid goes into the prompt itself.

What is new in Omni 1.1 Flash

  • Scene extension to 40 seconds. The model now reads the last 10 seconds of the existing video as context (earlier models looked at the final second only) and adds 10-second increments, so motion and story carry across the cut.
  • First and last frame control. Two images, two tags, and the model generates the continuous video between them - the cleanest way to land a camera move or a transition exactly where you planned it.
  • Video references. Up to three clips of up to three seconds each guide motion or look; their audio is ignored.
  • 360p drafts and 4K upscaling. Drafts render up to 60% faster and at a third of the cost of 720p; finals are 720p by default, 1080p or 4K upscaled.
SpecGemini Omni 1.1 Flash
Model IDgemini-omni-1.1-flash (Gemini API, Google AI Studio, Gemini Enterprise Agent Platform; Flow for AI Plus/Pro/Ultra; Gemini app)
Length3-10 s per call; scene extension in 10 s steps up to 40 s total, reading the last 10 s as context
Resolution360p draft; 720p default; 1080p and 4K upscaled
Aspect ratio / fps16:9 (default) or 9:16; 24 fps
Referencesimages via <IMAGE_REF_N>; up to three video clips of up to 3 s each via <VIDEO_REF_N> (their audio is ignored); <FIRST_FRAME> / <LAST_FRAME> for transitions
Audiogenerated by default; dialogue and sounds timed with [0-3s]-style timecodes; no audio uploads
Not supportedsystem instructions, temperature, top_p, stop sequences, negative prompts - write exclusions as plain constraints
ProvenanceSynthID watermark on every output

The prompt structure

Omni wants natural, literal instructions, not keyword lists. The same 2026 director's brief that works for Seedance 2.5 and Veo 3.1 works here, with two Omni-specific habits: tags for media and timecodes for time.

  1. Tags first. <FIRST_FRAME> and <LAST_FRAME> for a planned transition; <IMAGE_REF_N> with a role ("controls identity only"); <VIDEO_REF_N> for a motion or look you want copied.
  2. One sentence of intent - what the shot is, how long, in what space and light - followed by "keep everything else the same" when you edit or extend.
  3. Timecodes for every beat: [0-3s], [3-6s], one visible action per beat, and where the motion settles.
  4. Camera as one explicit move and lens.
  5. Audio described explicitly - room tone, effects, and dialogue in quotes with the second and the language.
  6. Constraints in plain words ("no on-screen text, no extra people") - negative prompts are not supported, so exclusions live in the prompt.

A complete example (8 s, first/last frame, dialogue)

<FIRST_FRAME> the model stands at the far end of a sunlit loft, back to camera, 16:9.
<LAST_FRAME> the same model in the same loft, now facing the camera at arm's length, soft smile.
<IMAGE_REF_1> controls her identity, face and body only - not her pose, clothing, background or lighting.

Generate an 8-second continuous shot between the first and last frame in the same loft and the same morning light. Keep everything else the same: the cream linen suit, the hair, the window light from the left.
[0-3s] she turns and walks toward the camera at an easy pace, heels on the wooden floor.
[3-6s] she slows, adjusts the lapel with her right hand, eyes on the lens.
[6-8s] she stops at arm's length and settles into the soft smile of the last frame.
Camera: slow push-in at eye level, 35 mm, no zoom, no cuts.
Audio: her heels on wood, quiet room tone, a distant tram outside; at 6 s she says "Ready when you are." in English, calm; no music.
Constraints: no on-screen text, no extra people, no watermark, no morphing.

Draft, extend, upscale

  1. Draft at 360p. Check timing, the settle at the end and the dialogue placement while it is cheap and fast.
  2. Extend in 10-second steps. Each extension reads the last 10 seconds, so write the next beat as a continuation ("she keeps walking, the camera keeps pushing in") and repeat the identity role and the continuity line verbatim. Stop at 40 seconds per scene.
  3. Upscale the approved take to 1080p or 4K. Upscaling sharpens; it does not fix a bad cut, so approve the motion at draft resolution first.

Editing existing footage follows the same rule: one plain instruction ("replace the red jacket with a cream linen suit, keep everything else the same") beats a paragraph of adjectives.

Omni 1.1 Flash versus the others

  • Choose Omni 1.1 Flash to extend a scene past 8 seconds inside Google's stack, to control the first and last frame, or to edit footage with a sentence.
  • Choose Veo 3.1 for 8-second photoreal hero clips with the most realistic light and up to three ingredient images.
  • Choose Seedance 2.5 for 30-second one-take stories with up to 50 references and explicit roles - see the Seedance 2.5 prompt guide.
  • Choose FLUX 3 Video for multilingual speech with lip sync in 5-20 second clips - see the FLUX 3 Video guide.

The full roster and prices are in AI video generators in 2026.

Common mistakes

  • Negative prompts or system instructions - unsupported; write constraints in the prompt.
  • Uploading an audio file as a reference - unsupported; describe the sound instead.
  • Adding new dialogue while extending a clip in which someone already speaks - not possible; plan lines per scene.
  • Video references longer than 3 seconds or more than three clips - trimmed or rejected.
  • Extending without repeating the identity role and continuity line - the character drifts at the cut.
  • Approving at 4K - draft at 360p first; upscaling does not repair motion.

FAQ

What is Gemini Omni 1.1 Flash?

Google's video generation model (model ID gemini-omni-1.1-flash), announced on August 27, 2026 as a production-ready update of Gemini Omni Flash. It generates and edits video with audio, and adds four controls: scene extension to 40 seconds with 10 seconds of prior context, first and last frame control, up to three 3-second video references, and 360p drafts that upscale to 1080p or 4K.

How long can a clip be?

Each call generates or continues 3 to 10 seconds. Scene extension adds 10-second increments and reads the last 10 seconds of the existing video as context (earlier models only looked at the final second), up to a total of 40 seconds per scene. Longer pieces are several scenes cut together.

How do first and last frame control work?

Attach two images and mark them in the prompt with the <FIRST_FRAME> and <LAST_FRAME> tags. The model generates continuous video between them, which is the cleanest way to get a planned camera move or a transition that lands exactly where you want it.

Does it generate audio and dialogue?

Yes. The model generates an audio track by default and you can time events and lines with timecodes such as [0-3s]. Uploading audio references is not supported, audio inside a video reference is ignored, and you cannot add new dialogue when extending an uploaded video in which someone is already speaking.

What is the cheapest way to iterate?

Draft at 360p, which Google says renders up to 60% faster and at a third of the cost of the standard 720p, fix the timing and the motion, then re-render the approved take at 720p and upscale to 1080p or 4K. Google publishes token-based pricing for the model; check its pricing table for current rates.

Omni 1.1 Flash or Veo 3.1 or Seedance 2.5?

Omni 1.1 Flash when you need to extend a scene past 8 seconds inside the Google ecosystem, control the first and last frame, or edit existing footage with a simple instruction. Veo 3.1 for 8-second photoreal hero clips with the most realistic light. Seedance 2.5 for 30-second one-take stories with up to 50 references and explicit roles. FLUX 3 Video for multilingual speech with lip sync in 5-20 second clips.


Want the tags, timecodes, camera, audio and constraints written for you? GoldenPrompts builds Omni, Veo, Seedance and Kling briefs from a few clicks. Free to start: 24 hours of everything, no card.