Black Forest Labs Enters Video Generation Race With FLUX 3, But Holds Back Key Details
Black Forest Labs has released FLUX 3, a multimodal AI model capable of generating images, video clips up to 20 seconds long, and audio from a single text prompt, marking the company's first public video generation tool. However, the limited early access rollout and absence of pricing information, service-level commitments, and complete benchmark details create uncertainty for enterprise buyers evaluating the model against competitors like Google's Gemini Omni Flash and Runway Gen-4.5.
How Does FLUX 3 Compare to Other Video Generation Models?
Black Forest Labs conducted preliminary preference testing on 10-second, 720p video clips with audio, comparing FLUX 3 against several established competitors. In these early evaluations, FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and Runway Gen-4.5 in 77% of comparisons. However, the results become more competitive when measured against cutting-edge models. FLUX 3 matched Google's Gemini Omni Flash at 52% preference, meaning the two models performed essentially equally in viewer evaluations.
The company also tested against Kling v3 Pro, Happy Horse v1, and Seedance 2.0, though these comparisons carry important caveats. The 60% preference over Kling v3 Pro and 52% tie with Seedance 2.0 represent meaningful data points, but the benchmark results are labeled as preliminary, describing a pre-release checkpoint rather than the final shipping model. This means actual customer performance may differ from these published figures.
What Makes FLUX 3's Architecture Different?
Rather than combining separate image, video, and audio models behind a common interface, Black Forest Labs trained FLUX 3 as a unified system across all three modalities simultaneously. The company calls this approach "visual intelligence," positioning it as a foundation for creative generation, simulation, computer use, and robotics applications. This joint training method differs fundamentally from how competitors assemble their tools.
"True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others," said Robin Rombach, co-founder and CEO of Black Forest Labs.
Robin Rombach, Co-founder and CEO at Black Forest Labs
The model builds on Self-Flow, Black Forest Labs' method for aligning multimodal understanding and generation within a single architecture, which the company publicized in March 2026. According to the company, significantly scaling up compute and data allowed training across video, images, and audio simultaneously, with testing showing that video generation and action prediction do not require separate foundations.
What Are the Availability and Pricing Details?
FLUX 3 is rolling out through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the upcoming open-source FLUX 3 Dev. Currently, FLUX 3 Video with optional native audio generation and FLUX 3 Action are entering a gated early access program that anyone can apply to, though Black Forest Labs must approve each application. FLUX 3 Image will roll out in the coming weeks, followed by general availability.
A significant limitation for developers and enterprises: FLUX 3 is not launching with downloadable weights or an open-source license. Black Forest Labs says faster and open-weight versions will arrive later in 2026, with FLUX 3 Dev offering "open-weight access to a multimodal backbone, for content creation and action prediction." However, this open-source variant arrives last in the rollout sequence, disappointing developers accustomed to receiving locally deployable versions alongside major announcements.
The company has not announced pricing, production service-level commitments, evaluation methodology, sample sizes, or rater counts. This absence of critical information prevents enterprise buyers from calculating total cost of ownership or independently reproducing the video quality comparisons. For context, competing models offer varying pricing structures: Google's Gemini Omni Flash costs $0.10 per second of generated 720p video, or roughly $1.00 for a 10-second clip, while Google's Veo 3.1 Fast tier costs $1.00 per 10-second 720p clip.
Steps to Evaluate FLUX 3 for Your Organization
- Request Early Access: Apply for FLUX 3's gated early access program to test video and action prediction capabilities before general availability, allowing hands-on evaluation of quality and performance.
- Compare Against Gemini Omni Flash: Since FLUX 3 and Google's Gemini Omni Flash performed equally in Black Forest Labs' preliminary testing, evaluate both models side-by-side once pricing and SLAs are published for FLUX 3.
- Wait for Pricing and Benchmarks: Hold off on major commitments until Black Forest Labs publishes complete pricing, service-level agreements, and final benchmark methodology during broader general availability.
- Consider Regional Constraints: If your organization operates in the European Economic Area, Switzerland, or the United Kingdom, note that Gemini Omni Flash currently restricts video editing for uploaded content in these regions, though FLUX 3's regional limitations remain unspecified.
What Gaps Remain Before Enterprise Adoption?
The missing details create friction for enterprise decision-making. Without published pricing, service-level commitments, and complete benchmark methodology, organizations cannot accurately compare FLUX 3 to alternatives or forecast implementation costs. The preliminary nature of the benchmark results, while not disqualifying, means the shipping model's actual performance against competitors remains unknown.
The delayed open-source release also affects adoption velocity. Open-weight FLUX models have driven significant community adoption and integration into third-party tools. Developers waiting for FLUX 3 Dev will have to rely on API access through Black Forest Labs or partners until the open-weight version arrives later in 2026, potentially slowing innovation and integration compared to competitors offering immediate local deployment options.
For enterprises evaluating video generation tools, the competitive landscape now includes FLUX 3's unified multimodal approach, Google's Gemini Omni Flash at known pricing and availability, and established players like Runway Gen-4.5 and Luma Ray 3.2. The choice will likely depend on specific use cases, regional requirements, and whether organizations prioritize immediate availability and transparent pricing over potential performance advantages that remain to be fully documented.