Blog 路 AI & Image Generation
A small German AI video model beats the big players
Olaf Lemmens, Founder NinA AI Agency 路 July 24, 2026 路 6 min read
Thursday is always the toughest day for AI releases. Models tumble over each other, everyone wants the attention, and usually half of them fade within a week. Last Thursday was one of those days. OpenAI launched a new voice for ChatGPT, Anthropic did something similar, and normally that would have been the talk of the day.
But my eye kept going back to something else. A small German company from Freiburg, Black Forest Labs, announced FLUX 3. A video model, I thought at first. Another one. Until I read on and got to the robots.
Because what makes this story remarkable is not the videos it makes. It is what happens when you put that same model on a factory floor at Audi.
TL;DR
Black Forest Labs, the German company behind the FLUX image models, launched FLUX 3: a model that learns image, video and audio in one architecture, with text-to-video up to 20 seconds and natively synchronised sound. The real surprise is that the exact same video backbone drives robots at Audi through ‘FLUX-mimic’, where they handle soft materials and cables that were impossible with conventional robotics. Some tasks can be fine-tuned with 30 minutes of robot data, and the system responds in roughly 101 milliseconds.
CEO Robin Rombach argues that content creation and physical AI draw on the same world knowledge. But almost nothing has been independently verified: no pricing, no parameter counts, no external benchmarks.
What FLUX was until now

A quick step back to who these people are, because it makes the story bigger. Black Forest Labs was founded in 2024 in Freiburg, by the researchers who earlier put latent diffusion and Stable Diffusion on the map. Robin Rombach is the CEO. If you have ever used a generative feature in Photoshop, Canva or Picsart, chances are there is a FLUX model underneath.
This is not a hobby club. In December 2025 they raised 300 million dollars in a Series B, at a valuation of 3.25 billion. Salesforce Ventures and Anjney Midha led that round, joined by a16z, NVIDIA, General Catalyst and Temasek. That announcement also revealed a previously undisclosed Series A. Added up, it comes to more than 450 million dollars.
That is the context FLUX 3 arrives in. Not as an experiment, but as the first multimodal foundation model from a company with serious money and serious pedigree.
Image, video and audio in one model

FLUX 3 launched on 23 July 2026. At its core is an architecture BFL calls ‘Self-Flow’: image, video and audio are learned jointly within one framework instead of as separate models you glue together.
The headline feature is text-to-video up to 20 seconds, with natively synchronised audio. So not making the video first and adding sound underneath afterwards, but dialogue, sound effects and ambient sound that come along immediately. The model can also continue from a starting frame, use videos as a reference, and string separate clips into longer scenes.
It was trained on tens of millions of hours of general video. And there is already a hint in there: the training set included hundreds of thousands of hours of video specifically about human and robot manipulation tasks. That last part turned out to be no side note.
Why a video model can suddenly drive a robot

BFL took the same video backbone and adapted it for robotics, under the name FLUX-mimic, developed with Swiss company mimic robotics. And they did not put it in a demo, but into production at Audi.
Christoph Schneider of Audi’s Production Lab says the robots solve work that would have been impossible with conventional robotics. Complex soft-body manipulation, he calls it. Think of gripping cables and handling deformable materials, exactly the kind of task where rigid pre-programmed robot arms get stuck. According to BFL, some tasks can be fine-tuned with just 30 minutes of robot data. And the system reportedly responds in roughly 101 milliseconds, close to human visual reflexes.
Rombach’s reasoning behind it is the actual claim. His position is that a model learning to generate convincing video must understand something about how the world works: how things move, fall, bend. And that this same world knowledge can drive a robot. Content creation and physical AI drawing from the same source. Four words: same model, different output.
And yet...

Let us stay sharp for a moment, because the story is less complete than the announcement suggests. BFL has released almost nothing you could use to check the claims. No pricing, no parameter count, no quantisations, no hardware requirements, no licence terms. And not a single image-model benchmark.
The video benchmarks that do exist come from BFL itself and have not been independently verified. In their own comparisons FLUX 3 beat Luma Ray 3.2 in 93 percent of cases and Runway Gen 4.5 in 77 percent. Against Seedance 2.0 and Gemini Omni Flash the outcomes were much closer. Those are interesting numbers, but they remain the seller’s numbers.
The rollout is also phased. FLUX 3 Video and FLUX 3 Action are now in early access through the API and private weights for a handful of partners. The image version follows in a few weeks, and an open-weight Dev version is planned for later in 2026. That approach, limited first and broader later, is one we now know from Anthropic and OpenAI. I would rather name what is actually available today than what might arrive later.
What this means for you in practice
Setting the hype aside, because the practical side matters more to most companies than the robot spectacle. Three things to take away.
1. If you already use FLUX through Canva, Photoshop or another platform, nothing changes for you in the short term. The open Dev version you could build with yourself only arrives later this year, and until then it runs behind APIs and private deals. Waiting is fine here.
2. Video with native audio in one step will save a lot of editing work over time. If you make content, keep an eye on this type of model, but do not count your money yet: without pricing, nobody knows what a minute of video will cost.
3. The robotics side is for now something for factories with Audi budgets, not for small and mid-sized businesses. But the underlying idea, that one model trained on the real world can do several tasks at once, is where the gains sit in the coming years. Not ten separate tools, but one foundation you point at different problems.
I opened this edition with a video model I almost skipped among all the Thursday releases. What makes it remarkable is not the video, but that the same technique teaches a robot at Audi to grip soft cables. Whether Rombach is right that generating images and driving robots really come from the same world knowledge, we will only know once the independent benchmarks are in. For now it is a fine promise with too few numbers behind it.
My question for you: do you see room for video models like this in your work, or does this stay something to watch from a distance for now? Let me know in the comments, I read all of them.
Until next time,
Olaf Lemmens
Founder NinA AI Agency
Have a look at www.nina-ai.nl