Black Forest Labs releases FLUX 3 video model with native audio, up to 20 seconds
On Aug 12 Black Forest Labs formally released the FLUX 3 video model via API, ranking #1 globally on text-to-video Elo, beating Gemini Omni Flash and Seedance 2.0.
On August 12, German AI startup Black Forest Labs (BFL) formally opened FLUX 3 video generation via its BFL API and select partners. FLUX 3 is a multimodal foundation model built on the Self-Flow unified architecture, integrating four specialized codecs for image, video, audio, and motion, with first-time support for native audio generation, capable of producing up to 20 seconds of HD synchronized audio-video content in a single shot.
FLUX 3 supports three audio modes (ambient, dialogue, music) with frame-accurate audio-video synchronization; the model releases in phases, with video available now and the image and open-source Flux3Dev versions upcoming. BFL says FLUX 3 supports text-to-video, image-to-video, keyframes, video continuation, and can include multiple scenes and multiple camera angles in a single output; the model can render typography directly in scenes, understand complex prompts, and leverage world knowledge for documentary-style content. FLUX 3 also supports lip-synced dialogue in over 14 languages.
BFL's own tests show FLUX 3 ranked #1 in text-to-video Elo at 1,135 points and image-to-video at 1,051 points; its results lead Gemini Omni Flash, Minimax H3, and Seedance 2.0. On a 720P resolution 10-second video benchmark, FLUX 3 defeated Luma Ray3.2 at 93% win rate and Runway Gen-4.5 at 77%.
To celebrate FLUX 3 video's debut at global rank #2, BFL launched limited free access until 11:59 PM PT on August 16, 2026; upcoming capabilities include 4K output, video editing, and multi-image/multi-video reference input support. Pricing is based on generated video seconds, with audio included.