Flux 3
One Model for Image, Video & Audio
Flux 3 is Black Forest Labs' new multimodal foundation model, generating image, video, and audio from a single unified architecture. Creations that are truer to life in every style.
FLUX 3 video generation is available on faktry: text-to-video, image-to-video, first/last-frame-to-video, and video extension, all with native audio.
Flux 3 in Action
See what Flux 3 can create, generated on faktry.
What Is FLUX 3?
FLUX 3 is a multimodal AI foundation model from Black Forest Labs, released on July 23, 2026. It generates image, video, and synchronized native audio from a single unified architecture built on Self-Flow, Black Forest Labs' method for aligning multimodal generation and understanding inside one model. FLUX 3 Video produces clips up to 20 seconds long in a single generation, with multilingual dialogue and sound-visual synchronization.
FLUX 3 at a glance
- Developer
- Black Forest Labs
- Released
- July 23, 2026
- Architecture
- Self-Flow multimodal foundation model
- Modalities
- Image, video, and native audio from one model
- Max video length
- 20 seconds per generation
- Audio
- Native, synchronized, multilingual dialogue
- Variants
- FLUX 3 Video, FLUX 3 Image, FLUX 3 Dev, FLUX 3 Action, FLUX-mimic
- Open weights
- Announced but not yet released. FLUX 3 Dev is confirmed as the open-weight variant, with no release date, license or Hugging Face repository published yet
- License
- Not published for FLUX 3. Earlier FLUX.1 [dev] weights ship under the BFL Non-Commercial License
- On faktry
- Live now: text-to-video, image-to-video, first/last-frame-to-video, and video extension
Last updated: · Model details published by Black Forest Labs
FLUX 3 vs Flux 2
What changes between Black Forest Labs' image-only Flux 2 and the multimodal FLUX 3.
| Capability | Flux 2 | FLUX 3 |
|---|---|---|
| Modalities | Image only | Image, video, and native audio |
| Video generation | Not supported | Up to 20 seconds per generation |
| Audio | Not supported | Native, synchronized, multilingual dialogue |
| Architecture | Image generation model | Self-Flow multimodal foundation model |
| Variants | Pro, Max, Flex | Video, Image, Dev, Action, FLUX-mimic |
| On faktry | Available now | Available now (video) |
FLUX 3 Video, Live on faktry
Four ways to generate and edit video with FLUX 3, each with native audio and billed per second of output.
Text to Video
Generate a video with synchronized native audio directly from a text prompt.
Image to Video
Animate a single starting image into a full video clip guided by your prompt.
First/Last Frame to Video
Pin a start frame and an end frame, and let FLUX 3 generate the motion in between.
Extend Video
Continue an existing clip from its final frames, keeping motion and style consistent.
Duration, resolution, aspect ratio, native audio, and safety tolerance are all configurable per generation.
What Makes Flux 3 Different
Built on Self-Flow technology, one model that understands how the world looks, moves, and sounds.
Truly Multimodal
A single model generates image, video, and audio, instead of separate systems bolted together.
Self-Flow Architecture
Flow matching aligns generation and understanding across data types, learning how objects hold together, how things move, and how events sound.
Video with Native Audio
Generates video with synchronized native audio up to 20 seconds, including text-to-video, image-to-video, and video-to-video.
Sharper Image Synthesis
Improved handling of complex prompts and text rendering compared to earlier FLUX models, across styles, aspect ratios, and resolutions.
Multilingual by Design
High-accuracy multilingual text rendering in images and multilingual dialogue in generated video.
Lifelike Motion & Sound
Strong performance on facial expressions and sound-visual synchronization, with keyframe-to-video control.
Capabilities
A closer look at what Flux 3 can generate, across video and image.
Video Generation
Image Synthesis
Early Benchmarks
In Black Forest Labs' preliminary text-to-video evaluations, FLUX 3 Video was preferred over other leading video models in head-to-head human comparisons.
Preliminary results published by Black Forest Labs at announcement. Not independently verified by faktry.
Where Flux 3 Will Fit In
One model covering the workflows that used to need several separate tools.
Content Creation
Go from a single prompt to on-brand image and video content, in one consistent style.
Marketing Video
Produce short video with native audio and multilingual dialogue for campaigns and social.
Product Visualization
Turn product concepts into photorealistic images and motion without a separate render pipeline.
Global Content
Accurate multilingual text in images and multilingual dialogue in video for international audiences.
Can I try Flux 3 for free?
Yes. Every new account gets 100 credits, no card required, and FLUX 3 video generation is unlocked from the start.
Free credits to try us
Frequently Asked Questions
What is Flux 3?
Flux 3 is Black Forest Labs' new multimodal foundation model. It generates image, video with native audio, and text from one unified model built on Self-Flow technology.
Can I try Flux 3 for free?
Yes. Sign up at faktry.ai and you get 100 free credits instantly, no credit card required. That is enough to generate your first FLUX 3 videos straight away. After that there is no subscription: FLUX 3 is billed per second of output, so you pay only for what you actually generate.
When was Flux 3 released?
Black Forest Labs released FLUX 3 on July 23, 2026, with video generation available first. faktry brought FLUX 3 video online shortly after. Image capabilities and the wider lineup (FLUX 3 Image, FLUX 3 Dev, FLUX 3 Action, and FLUX-mimic) are rolling out over the following months.
Are the Flux 3 weights available on Hugging Face?
Not yet. Black Forest Labs has confirmed FLUX 3 Dev as an open-weight release, but as of August 2026 there is no published release date, no license, no disclosed parameter count and no FLUX 3 repository on Hugging Face. The Black Forest Labs organisation page on Hugging Face is where the weights will appear when they land. In the meantime FLUX 3 is reachable through hosted APIs, including faktry.
Can I use Flux 3 on faktry today?
Yes. FLUX 3 video generation is live on faktry now: text-to-video, image-to-video, first/last-frame-to-video, and video extension, all with native audio and billed per second of output. Flux 2, Sora 2, Veo 3.1, Kling 3.0, and Seedance are also available.
How is Flux 3 different from Flux 2?
Flux 2 is an image-only model. Flux 3 is multimodal: it generates image and video with native audio from a single model, with improved prompt handling and multilingual text accuracy.
What can Flux 3 actually generate?
Video up to 20 seconds with native audio and multilingual dialogue, plus images across a wide range of styles, aspect ratios, and resolutions, all from the same underlying model.
Which other models can I use on faktry right now?
Flux 2, Sora 2, Veo 3.1, Kling 3.0, Seedance, and more are already live on faktry alongside FLUX 3. Browse the full lineup on our AI Models page.