FLUX 3 wants to be one model for images, video, audio, and even robot actions
Germany's Black Forest Labs unveiled FLUX 3, a single model trained to generate images and short video with sound, and to extend to robot vision. Access is limited to start.
Black Forest Labs, the Freiburg-based lab behind the popular FLUX image models, announced FLUX 3 on July 23. The pitch is ambitious: one model, trained across several kinds of media at once, that can generate images, produce video clips with native sound up to 20 seconds long, edit pictures, render readable text, and even predict actions for robots. “Multimodal” is the key word here, and it means the model handles more than one type of data (images, video, audio) rather than just text.
The technical claim worth understanding is how it is built. Many “all-in-one” tools are really several separate models stitched together behind one interface. Black Forest Labs says FLUX 3 is instead jointly trained across images, video, and audio, sharing a single set of weights. Weights are the internal numbers a model learns during training; sharing one set means the same underlying system does all the jobs, rather than handing off between specialists. In theory that makes results more consistent across formats.
What is behind this is a land grab in creative AI. For a couple of years, image generation and video generation lived in different tools from different companies. The frontier now is a single model that does the lot, and a European lab planting a flag there matters for anyone who cares about not depending only on US providers. There is also a longer game: extending the same model to robot vision and actions points at “physical AI,” where the software that dreams up a picture and the software that steers a machine start to converge.
Some grounded caveats. FLUX 3 is rolling out in stages, not all at once. FLUX 3 Video is in early access now; FLUX 3 Image and FLUX 3 Action follow in the coming weeks. An open-weight version called FLUX 3 Dev, which you could run and customize on your own hardware, is planned for later in 2026 but is not here yet. So the headline capabilities are real but gated, and the version most hobbyists will want is still a promise.
What this means for you: If you make things, images, short videos, social clips, this is worth watching, because one tool that does image, video, and sound could replace a small stack of separate apps. If you are just curious, the more interesting signal is direction: AI media tools are collapsing into single models, and one of the leaders is European. For most of us, the practical moment arrives when FLUX 3 Dev ships as open weights and you can try it without a queue. Until then, keep expectations grounded and treat the demos as demos.
Sources
ChatGPT can now read your health records, which is handy and worth thinking about
OpenAI launched Health in ChatGPT for US adults on July 23, letting you connect medical records and Apple Health. It is genuinely useful, but privacy experts flag a real trade-off.