Black Forest Labs' FLUX 3 Ditches Still Images for Video,And Now Controls Robot Hands
Black Forest Labs unveiled FLUX 3 on Thursday, marking the company's first leap from still-image generation into video creation, with the model also powering a new robotics system for manufacturing. The German AI lab trained the new system on images, video, and audio simultaneously within a single shared model, a technique known as multimodality that allows one AI system to learn and generate multiple types of information together rather than relying on separate tools bolted side by side.
The video capabilities represent a significant departure from FLUX's image-focused legacy. FLUX 3 produces video clips up to 20 seconds long with audio generated alongside the picture and synchronized to on-screen action, including dialogue, sound effects, and ambient noise. In early human evaluations, reviewers preferred FLUX 3's output over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93% of matchups.
How Does FLUX 3 Compare to Other AI Video Models?
FLUX 3 performs competitively across the broader AI video landscape. The model beat Gemini Omni and Seedance in 52% of evaluations, though these preference tests reflect human judgment rather than fixed scoring metrics. Beyond video, FLUX 3 maintains strong capabilities for still-image generation, demonstrating versatility across photorealistic and stylized outputs.
- Runway Gen-4.5: FLUX 3 outperformed this competitor in 77% of direct comparisons.
- Luma Ray 3.2: FLUX 3 won 93% of matchups against this model.
- Gemini Omni and Seedance: FLUX 3 beat these models in 52% of evaluations, indicating closer competition.
Why Does Video Prediction Matter for Robotics?
Black Forest Labs frames FLUX 3 as more than a content creation tool. The company's strategic bet is that learning to predict video requires understanding the physics underlying motion, including weight, contact, and timing, which are precisely the capabilities needed for machines to move through the physical world.
This insight led to FLUX-mimic, a robotics system developed in partnership with Zurich-based mimic robotics. The system takes FLUX 3's video-prediction engine and adds a lightweight decoder, a small add-on component that translates the model's internal understanding of motion into actual robot movements. Car manufacturer Audi is already testing FLUX-mimic on tasks like fitting flexible door seals, work that conventional automation has historically struggled to handle.
"A model that only learns images can only generate images," said Robin Rombach, co-founder and CEO of Black Forest Labs.
Robin Rombach, Co-founder and CEO at Black Forest Labs
Audi's testing demonstrates real-world manufacturing applications. The company reports that robots using FLUX-mimic now "solve complex soft-body manipulation work" that older machines could not perform. The full system reacts in approximately 101 milliseconds, a response time in the neighborhood of human visual reflexes.
"Audi represents the kind of manufacturing partner we built FLUX-mimic for," said Stephan-Daniel Gravert, co-founder of mimic robotics.
Stephan-Daniel Gravert, Co-founder at mimic robotics
How Did Black Forest Labs Overtake Stability AI in Image Generation?
FLUX 3's emergence reflects a dramatic shift in the AI image-generation landscape. Black Forest Labs was founded in August 2024 by veteran researchers who had previously helped build the original Stable Diffusion models at Stability AI. The company's FLUX models quickly outclassed Stability's own offerings, including the underwhelming Stable Diffusion 3.
The open-source FLUX Dev and Schnell models captured the "best open-source image generator" title that AI artists had expected Stable Diffusion 3.5, Stability's do-over release, to eventually claim. That title never materialized for Stability. FLUX 1.1 Pro went on to top the Artificial Analysis image arena by October 2025, though that version was not open-source. FLUX 2, released in November 2025, did not achieve the same popularity as the original FLUX models.
The open-source crown held by FLUX lasted until Alibaba's Z-Image Turbo dethroned it in late 2025, matching FLUX's quality on lower-end consumer graphics cards. One user on CivitAI, a platform for sharing AI models, remarked that Z-Image Turbo was "what SD3 was supposed to be," underscoring Stability AI's failure to deliver a competitive successor to Stable Diffusion.
When Will FLUX 3 Be Available?
FLUX 3 is not yet fully open to the public. Video and Action capabilities are currently in early access through application programming interfaces (APIs) and select partners, with mimic robotics among them. Image generation will follow "in the coming weeks," according to Black Forest Labs. The open-weight Dev version, the only tier BFL plans to release for local use on personal computers, is not expected until later in 2026.
This phased rollout strategy contrasts with the company's earlier approach to FLUX models, which were released more openly. The delayed availability of the open-weight version reflects the competitive pressures and commercial opportunities surrounding advanced AI video generation, a space increasingly dominated by well-funded startups and tech giants.