Published 2026-09-15 · Updated 2026-09-15

FLUX 3 Video Prompt Guide: 20-Second Clips with Native Audio

Short answer: FLUX 3 Video renders picture and sound together, so the prompt has to direct both. Write a FORMAT line, give the keyframes their timestamps, describe the STARTING STATE, lay out a TIMELINE with one visible action per beat, fix the CAMERA and the CONTINUITY, and write the AUDIO as a scored line — who speaks, in which language, at which second, plus the effects and the room tone. Render a $0.06 per second draft first, then enhance it.

What FLUX 3 Video is

FLUX 3 is Black Forest Labs' multimodal model — one backbone for images, video and audio — announced on July 23, 2026 in early access. Its video mode reached the BFL API and select partners on August 4, 2026. The headline is not the length but the sound: speech in a dozen-plus languages with lip sync, effects and ambience are created with the frames, not bolted on afterwards.

SpecFLUX 3 Video
Length5-20 s (video continuation 5-15 s), 24 fps
ResolutionHD 1920 x 1088 for 16:9; FHD through the built-in upsampler
Aspect ratios21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16
Audioon by default: multilingual speech with lip sync, effects, ambience; can be switched off
Inputstext; 1-10 images as start frame, start+end pair or timestamped keyframes; up to 4 s of existing video and audio for continuation
Price (BFL, Sept 2026)$0.17/s HD, $0.29/s FHD; draft $0.06/s; continuation $0.43/s HD, $0.54/s FHD
AvailabilityBFL API and dashboard since Aug 4, 2026; partner platforms; no open weights

The brief structure

The same 2026 director's-brief pattern that works for Seedance 2.5 and Kling 3.0 works here, with two FLUX-specific blocks: KEYFRAMES with timestamps instead of reference roles, and an AUDIO line that is scored to the second.

  1. FORMAT — length, aspect ratio, hd or fhd, one take or shots, audio on or off.
  2. KEYFRAMES — which image sits at which second and what must stay identical between them.
  3. STARTING STATE — subject, camera and light at 0 s.
  4. TIMELINE — second ranges, one visible action per beat, where the last motion settles.
  5. CAMERA — one move, lens, what is forbidden.
  6. CONTINUITY — invariants: product, face, clothes, light.
  7. AUDIO — effects with timestamps, room tone, dialogue in quotes with speaker, language and second, music or explicitly none.
  8. CONSTRAINTS — what must not happen.

A complete example (12 s, 9:16, product with a spoken line)

FORMAT: 12 s, 9:16, fhd, one continuous take, audio on.
KEYFRAMES: [0, image1] the glass bottle on the marble counter, cap on, morning light from the left; [8, image2] the same bottle with the cap off and one drop on the rim. Keep the label, the cap and the marble identical between the two.
STARTING STATE: the bottle stands centred, still; a hand is out of frame to the right.
TIMELINE:
0-4 s - the light slides slowly across the marble; a hand enters from the right and lifts the cap.
4-8 s - close on the rim: one drop forms and falls in real time; the hand sets the cap down out of frame.
8-12 s - the camera eases back to a hero angle and holds; the bottle settles centred for the last second.
CAMERA: locked-off macro at 100 mm for the first eight seconds, then a slow pull-back; no whip, no zoom.
CONTINUITY: same bottle, label and lighting throughout; the hand has short nails and no rings.
AUDIO: a soft cap click at 3.5 s, a single drop at 6 s, quiet room tone throughout; at 9 s a calm female voice (English, British) says "Made in one drop."; no music.
CONSTRAINTS: no on-screen text, no second product, no logo changes, no slow motion, no watermark.

Writing the audio line

  • Dialogue in straight quotes, with the speaker, the language and the second: at 9 s a calm female voice (English, British) says "Made in one drop."
  • One short line per clip under 15 s. Speech needs about three seconds per sentence plus silence around it; a 12-second clip carries one line, a 20-second clip two or three.
  • Effects with timestamps so they land on the action: "cap click at 3.5 s, drop at 6 s".
  • Name the room tone ("quiet room tone", "rain on the window") — silence without it is filled with a random bed.
  • Say "no music" when you mean it; otherwise the model may score the clip.
  • Speaker on screen must be visible when the line plays if you want lip sync; a voice-over is fine when nobody is on screen.

Keyframes with timestamps

Image-to-video accepts one to ten images. A single image is the start frame; two are a start-and-end pair; more become timestamped keyframes, written as [0, image1], [8, image2]. The model interpolates the motion between them, so the images must agree on everything you do not want to change — product, label, lighting direction, framing — and differ only in the state you want animated. Two product photos taken from the same tripod position are ideal; two photos from different angles produce a morph.

Draft first, then enhance

Every FLUX 3 Video request accepts draft: true: a fast 720p preview at $0.06 per second instead of $0.17. Check the timing of the line, the settle at the end and the camera move; fix the brief if needed; then resubmit the returned draft_cache with mode: "draft_enhance" to render the approved draft at full quality. A 12-second clip costs $0.72 to draft and about $2 to finish in HD — far cheaper than three blind full renders.

Where it beats the others, and where it does not

  • Choose FLUX 3 Video for 10-20 second product, brand and explainer clips with spoken lines in many languages, keyframe control and the draft-enhance loop.
  • Choose Veo 3.1 for 8-second photoreal hero shots with native audio; the most realistic light in 2026.
  • Choose Seedance 2.5 for 30-second one-take stories driven by up to 50 references with explicit roles — see the Seedance 2.5 prompt guide.
  • Choose Kling 3.0 for cut-based scenes with named speakers and per-shot durations.

Migrating from Sora? The prompt structure above maps one-to-one; the Sora shutdown migration guide has the feature-by-feature table.

Common mistakes

  • Dialogue written as description ("she talks about the product") — no line is spoken.
  • Two lines in an 8-second clip — the second one is rushed or cut.
  • Keyframes from different angles — the model morphs instead of animating.
  • Combining image and video inputs in one request — not supported; pick one.
  • Skipping the draft — full renders at $0.17-0.29 per second add up fast.
  • Spec tokens ("4K 60fps") and negative prompts — use FORMAT and CONSTRAINTS instead.

FAQ

What is FLUX 3 Video?

The video mode of FLUX 3, Black Forest Labs' multimodal model announced on July 23, 2026. FLUX 3 Video became available through the BFL API and select partners on August 4, 2026. It renders 5 to 20 seconds at 24 fps in HD (1920 x 1088 for 16:9) with Full HD via the built-in upsampler, and generates audio alongside the frames: multilingual speech with lip sync, sound effects and ambience.

Does it really generate dialogue and lip sync?

Yes. Audio is on by default (generate_audio: true) and the model renders speech in English dialects, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi and more, with lip sync, plus effects and ambient sound created with the frames. Write the line in quotes, say who speaks, in which language and at what second.

What inputs does it take?

Text only; text plus 1 to 10 images used as a single start frame, a start-and-end pair, or timestamped keyframes such as [0, image], [2.5, image]; or an existing clip as start_video for continuation, which carries up to four seconds of its video and audio forward. Image and video inputs cannot be combined in one request.

What does it cost?

BFL list prices in September 2026: text-to-video and image-to-video $0.17 per second in HD and $0.29 per second in FHD, draft mode $0.06 per second; video continuation $0.43 per second HD, $0.54 FHD, draft $0.12. A 12-second HD clip is about $2, its draft $0.72. Partner platforms set their own rates.

What is draft mode?

A fast, cheaper 720p preview of the same prompt. Set draft: true, check timing, dialogue and camera, then resubmit the returned draft_cache with mode "draft_enhance" to render the approved draft at full quality — the enhanced render follows the draft's motion instead of rolling new dice.

FLUX 3 Video or Veo 3.1 or Seedance 2.5?

FLUX 3 Video for 10-20 second clips that need spoken lines in many languages with lip sync and a draft-then-enhance workflow; Veo 3.1 for 8-second realism with native audio and up to three ingredient images; Seedance 2.5 for 30-second one-take stories with up to 50 references and explicit reference roles. Kling 3.0 remains the choice for cut-based scenes with named speakers.


Want the brief written for you — keyframes, timeline, camera, continuity, a scored audio line and constraints? GoldenPrompts builds FLUX 3 Video, Veo, Seedance and Kling briefs from a few clicks. Free to start: 24 hours of everything, no card.