Black Forest Labs launches FLUX 3 multimodal AI model with video generation and robotics capabilities

2 weeks ago 10



Black Forest Labs just dropped FLUX 3, a multimodal AI model that can generate 20-second video clips with synchronized audio, process images, and predict actions for robotics. The model launched with early access on July 23, representing BFL’s first truly integrated foundation model. FLUX 3 operates across images, video, audio, and physical action prediction within a single architecture. What FLUX 3 actually does The headline feature is text-to-video generation that produces clips up to 20 seconds long, complete with native audio that stays in sync with the visuals. That means dialogue, sound effects, and ambient noise all align with what’s happening on screen, rather than requiring separate audio generation and manual stitching. BFL calls the underlying technology its “Self-Flow” approach, which enables aligned understanding and generation of multimodal data within one unified system. Early evaluations suggest the model is performing well against its peers. FLUX 3 was preferred over competing models like Runway Gen-4.5 in 77% of comparisons, according to initial benchmarks. The phased rollout plan promises to expand access across Video/Audio, Image, and Action functionalities over...

Read Entire Article