Tech

BFL launches FLUX 3, a unified multimodal AI model for video, audio, and image generation

BFL has released FLUX 3, a multimodal frontier model available in Early Access that generates videos with native audio up to 20 seconds long. Preliminary tests show the model surpassing rivals including Runway Gen-4.5 and Luma Ray 3.2.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
New foundation model integrates spatial, temporal, and acoustic data to outperform key competitors in early evaluations

BFL has launched FLUX 3, a new multimodal foundation model designed to jointly learn from images, video, and audio within a single unified architecture. Described by the company as a "multimodal frontier model," FLUX 3 integrates spatial, temporal, and acoustic data to build a cohesive representation of the world, moving beyond the isolation of individual sensory inputs. The model is now available in Early Access, allowing developers to generate diverse videos with native audio up to 20 seconds in length, as well as synthesise and edit images.

The architecture utilises a "Self-Flow" approach to efficiently align multimodal generation and understanding. By scaling up compute and data resources, BFL trained FLUX 3 to process these modalities simultaneously, arguing that mutual constraints between sound, motion, and visuals provide a more accurate depiction of underlying reality than any single modality alone. This allows the model to mix inputs, generating video and audio jointly from text prompts or using visual references to maintain consistency across sequences.

In preliminary evaluations, FLUX 3 demonstrated strong performance against established competitors. The model was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons. It also outperformed Grok Imagine Video in up to 69% of tests, Kling v3 Pro in 60%, and Gemini Omni Flash in 52%. Early results highlight the model’s ability to capture human facial expressions, associate sounds with physical events, and render high-accuracy text in multiple languages.

Beyond content creation, BFL is extending FLUX 3’s capabilities into physical AI. The company is developing FLUX-mimic, a video-action model for robotics, in partnership with Mimic Robotics. Mimic was among the first partners to gain early access to test the model’s application in dexterous manipulation and production deployment. BFL notes that the pretrained video backbone can serve as a dynamics-aware foundation, allowing specialized action models to be fine-tuned with limited task-specific data.

Looking ahead, BFL plans to open early access for FLUX 3 Image in the coming weeks, following improvements in handling complex prompts and text generation. The company is also hiring in Germany and the US to support its mission to unify perceptual, action, and language prediction in future models. As the model continues to evolve through its early access phase, BFL expects further refinements in safety, consistency, and generation quality.

Continue reading

More from Tech

Read next: Baidu and Lyft begin London robotaxi trials with Freenow
Read next: Nonprofits brace for AI IPO wealth wave as OpenAI and Anthropic prepare for public listing
Read next: Report finds Hugging Face lacks safeguards against nonconsensual deepfake generation