Black Forest Labs has officially launched FLUX 3, a groundbreaking artificial intelligence model that marks a significant departure from its predecessors. For the first time, the German AI lab’s flagship product is capable of generating not only still images but also dynamic video content, complete with synchronized audio. This advancement is a direct result of a novel training approach where the system processed images, video, and audio simultaneously within a single, unified architecture. This "multimodality" allows FLUX 3 to develop a deeper, more integrated understanding of information, moving beyond the limitations of separate, bolted-on tools.
The introduction of video generation capabilities is the most striking feature of FLUX 3. The model can produce video clips of up to 20 seconds in length, with accompanying audio meticulously synced to the visual action. This includes dialogue, sound effects, and ambient noises, creating a more immersive and realistic output. Early evaluations have positioned FLUX 3 as a strong contender in the rapidly evolving AI video generation landscape. In head-to-head comparisons, human reviewers reportedly favored FLUX 3’s output over Runway Gen-4.5 in 77% of instances and over Luma Ray 3.2 in a remarkable 93% of comparisons. While slightly trailing behind Gemini Omni and Seedance in some aspects, FLUX 3 still outperformed these models in 52% of evaluations, underscoring its competitive edge.
The implications of this multimodal approach extend far beyond content creation. Black Forest Labs views FLUX 3 as a foundational technology with the potential to bridge the gap between digital intelligence and physical action. "A model that only learns images can only generate images," stated Robin Rombach, co-founder and CEO of Black Forest Labs. The company’s strategic bet is that by learning to predict video, the AI also implicitly learns the fundamental physics governing the real world – concepts such as weight, contact, and timing. This acquired understanding is precisely what is needed for machines to interact effectively and intelligently within the physical environment.

FLUX-mimic: Bridging the Digital and Physical Divide
This vision is being actively realized through FLUX-mimic, a project developed in collaboration with Zurich-based mimic robotics. FLUX-mimic leverages FLUX 3’s sophisticated video-prediction engine and augments it with a lightweight "decoder." This add-on component translates the AI’s internal understanding of motion and interaction into actionable commands for robotic systems. The potential applications are vast, particularly in industries requiring intricate manipulation and adaptability.
Automaker Audi is already exploring the capabilities of FLUX-mimic. The company is reportedly testing the system on complex tasks such as fitting flexible door seals, a job that has historically presented significant challenges for conventional automation. The ability of FLUX-mimic to handle such delicate and nuanced operations highlights the practical value of AI models that possess a grounded understanding of physical dynamics. According to Stephan-Daniel Gravert, co-founder of mimic robotics, "Audi represents the kind of manufacturing partner we built FLUX-mimic for." Christoph Schneider of Audi further elaborated on the benefits, stating that the robots equipped with this technology can now "solve complex soft-body manipulation work" that was previously beyond the reach of older robotic systems. The responsiveness of the full FLUX-mimic system is also impressive, with reaction times reportedly around 101 milliseconds, a speed comparable to human visual reflexes, suggesting a high degree of real-time operational capability.
A Legacy of Innovation and Disruption
Black Forest Labs’ journey to FLUX 3 is rooted in a history of significant contributions to the AI landscape, particularly in image generation. The company was founded in August 2024 by veteran researchers who played instrumental roles in the development of the original Stable Diffusion models at Stability AI. Shortly after its inception, Black Forest Labs released its initial FLUX models, which quickly established themselves as formidable competitors, outperforming established players like MidJourney and even surpassing Stability AI’s own Stable Diffusion 3.
The early open-source releases, FLUX Dev and Schnell, captured the attention of the AI art community, filling a void that many felt was not adequately addressed by Stability AI’s subsequent efforts. These models were widely lauded as the "best open-source image generator," a title that AI artists had anticipated Stable Diffusion 3.5, Stability’s later iteration, would eventually reclaim. However, this did not materialize as expected. In October of the same year, FLUX 1.1 Pro emerged as the leader in the Artificial Analysis image arena, though it was not an open-source offering.

The trajectory continued with the release of FLUX.2 in November 2025. While this version also saw development, it did not achieve the same level of widespread popularity as its predecessors. The open-source leadership held by the initial FLUX models was eventually challenged and surpassed in late 2025 by Alibaba’s Z-Image Turbo. This new model achieved comparable quality on lower-end consumer graphics cards, prompting user comments on platforms like CivitAI that "This is what SD3 was supposed to be."
FLUX 3’s Strategic Rollout and Future Prospects
FLUX 3 represents Black Forest Labs’ strategic comeback in the competitive AI market. Currently, the full capabilities of FLUX 3, including video and action generation, are in early access through APIs and select partners, with mimic robotics being a key collaborator. The still image generation functionalities are slated for release "in the coming weeks," according to the company. Black Forest Labs intends to release an open-weight "Dev" version, which will be the only tier available for local use, later in 2026. This tiered release strategy suggests a focus on controlled deployment and partnership development for the more advanced multimodal features, while gradually making certain aspects of the technology more broadly accessible.
The company’s emphasis on building models that understand not just visual data but also the underlying physics and temporal dynamics of the world signals a broader ambition. By integrating audio and video processing into a single, cohesive AI, Black Forest Labs is not just aiming to create more sophisticated generative tools but also to lay the groundwork for AI systems that can more seamlessly interact with and understand the complexities of the physical universe. The partnership with mimic robotics and the early adoption by industry leaders like Audi are strong indicators that FLUX 3’s impact could extend far beyond the realm of digital content creation into tangible, real-world applications. The company’s historical success in disrupting the image generation market suggests that FLUX 3 could similarly redefine the landscape of multimodal AI and robotics.
