10 hours ago
- FLUX 3 is a new multimodal foundation model that jointly learns from images, videos, and audio within a unified architecture, aiming to learn a representation of the real world.
- The model is built on the principle that no single modality provides a complete description; instead, multiple modalities provide mutual constraints that reveal underlying reality.
- FLUX 3 can generate and edit images, videos with audio, and supports capabilities like text-to-video, image-to-video, video-to-video, multilingual dialogue, and action prediction.
- Preliminary evaluations show FLUX 3 outperforms several other models in video generation, with up to 93% preference over some competitors.
- The launch plan includes making FLUX 3 available through APIs, private weight access, and eventually open-weight access for multimodal backbone use.
- Future goals include unifying perceptual, action, and language prediction into a single model.