BFL launches FLUX 3, a unified multimodal AI model for video, audio, and image generation
BFL has released FLUX 3, a multimodal frontier model available in Early Access that generates videos with native audio up to 20 seconds long. Preliminary tests show the model surpassing rivals including Runway Gen-4.5 and Luma Ray 3.2.
BFL has launched FLUX 3, a new multimodal foundation model designed to jointly learn from images, video, and audio within a single unified architecture. Described by the company as a "multimodal frontier model," FLUX 3 integrates spatial, temporal, and acoustic data to build a cohesive representation of the world, moving beyond the isolation of individual sensory inputs. The model is now available in Early Access, allowing developers to generate diverse videos with native audio up to 20 seconds in length, as well as synthesise and edit images.
The architecture utilises a "Self-Flow" approach to efficiently align multimodal generation and understanding. By scaling up compute and data resources, BFL trained FLUX 3 to process these modalities simultaneously, arguing that mutual constraints between sound, motion, and visuals provide a more accurate depiction of underlying reality than any single modality alone. This allows the model to mix inputs, generating video and audio jointly from text prompts or using visual references to maintain consistency across sequences.
In preliminary evaluations, FLUX 3 demonstrated strong performance against established competitors. The model was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons. It also outperformed Grok Imagine Video in up to 69% of tests, Kling v3 Pro in 60%, and Gemini Omni Flash in 52%. Early results highlight the model’s ability to capture human facial expressions, associate sounds with physical events, and render high-accuracy text in multiple languages.
Beyond content creation, BFL is extending FLUX 3’s capabilities into physical AI. The company is developing FLUX-mimic, a video-action model for robotics, in partnership with Mimic Robotics. Mimic was among the first partners to gain early access to test the model’s application in dexterous manipulation and production deployment. BFL notes that the pretrained video backbone can serve as a dynamics-aware foundation, allowing specialized action models to be fine-tuned with limited task-specific data.
Looking ahead, BFL plans to open early access for FLUX 3 Image in the coming weeks, following improvements in handling complex prompts and text generation. The company is also hiring in Germany and the US to support its mission to unify perceptual, action, and language prediction in future models. As the model continues to evolve through its early access phase, BFL expects further refinements in safety, consistency, and generation quality.

