Flux 3
One Model for Image, Video & Audio
Flux 3 is Black Forest Labs' new multimodal foundation model, generating image, video, and audio from a single unified architecture. Creations that are truer to life in every style.
Flux 3 was just announced by Black Forest Labs. We'll switch this on for faktry the moment we get access.
What Makes Flux 3 Different
Built on Self-Flow technology, one model that understands how the world looks, moves, and sounds.
Truly Multimodal
A single model generates image, video, and audio, instead of separate systems bolted together.
Self-Flow Architecture
Flow matching aligns generation and understanding across data types, learning how objects hold together, how things move, and how events sound.
Video with Native Audio
Generates video with synchronized native audio up to 20 seconds, including text-to-video, image-to-video, and video-to-video.
Sharper Image Synthesis
Improved handling of complex prompts and text rendering compared to earlier FLUX models, across styles, aspect ratios, and resolutions.
Multilingual by Design
High-accuracy multilingual text rendering in images and multilingual dialogue in generated video.
Lifelike Motion & Sound
Strong performance on facial expressions and sound-visual synchronization, with keyframe-to-video control.
Capabilities
A closer look at what Flux 3 can generate, across video and image.
Video Generation
Image Synthesis
Early Benchmarks
In Black Forest Labs' preliminary evaluations, FLUX 3 Video was preferred over other leading video models.
Preliminary results published by Black Forest Labs at announcement. Not independently verified by faktry.
A First Look
Early samples generated with FLUX 3, shared by Black Forest Labs.
Where Flux 3 Will Fit In
One model covering the workflows that used to need several separate tools.
Content Creation
Go from a single prompt to on-brand image and video content, in one consistent style.
Marketing Video
Produce short video with native audio and multilingual dialogue for campaigns and social.
Product Visualization
Turn product concepts into photorealistic images and motion without a separate render pipeline.
Global Content
Accurate multilingual text in images and multilingual dialogue in video for international audiences.
Frequently Asked Questions
What is Flux 3?
Flux 3 is Black Forest Labs' new multimodal foundation model. It generates image, video with native audio, and text from one unified model built on Self-Flow technology.
Can I use Flux 3 on faktry today?
Not yet, Flux 3 was just announced by Black Forest Labs. We're bringing it to faktry as soon as it becomes available to us.
How is Flux 3 different from Flux 2?
Flux 2 is an image-only model. Flux 3 is multimodal: it generates image and video with native audio from a single model, with improved prompt handling and multilingual text accuracy.
What can Flux 3 actually generate?
Video up to 20 seconds with native audio and multilingual dialogue, plus images across a wide range of styles, aspect ratios, and resolutions, all from the same underlying model.
Which other models can I use on faktry right now?
Flux 2, Sora 2, Veo 3.1, Kling 3.0, Seedance, and more are already live on faktry. Browse the full lineup on our AI Models page.