// announcement 002 / models

FLUX 3 is live on Skiprr

Black Forest Labs built one model that learns from images, video and audio together. It generates video with the sound already in it, up to 20 seconds a pass, from keyframes you set yourself. It is in the model picker on every Skiprr video production today.

ModelsAug 11, 2026

Black Forest Labs released FLUX 3, the first model they have trained entirely across images, video and audio at once, and Skiprr picked it up in early access. Open any video production and it is there in the model picker. No keys, no waitlist, no separate account, billed in the same credits as everything else in your workspace.

one prompt, one pass, picture and sound generated togethertext to video

The argument behind the model is worth understanding, because it explains what the output feels like. Images capture space at an instant. Video adds time, and with it the physics. Audio carries the causes that vision alone misses, the impact you hear before you see what caused it. Train on any one and you learn that projection well. Train on all three at once and they constrain each other: the sound has to match the impact, the motion has to obey the mass. FLUX 3 is the first model Black Forest Labs have built entirely on that principle, and the joins are where you notice it.

The sound is generated with the picture, not after it

Audio is not a pass bolted on at the end. Ambience, effects and speech come out of the same generation as the frames, which is why footsteps land on the footfall and a voice belongs to the face saying it. Dialogue goes in double quotes and comes back lip synced, in a range of languages. Ask for a silent clip and you get one, but silence is now a decision you make rather than the default you work around.

dialogue, lip movement and voice produced in the same passlipsync, multilingual
room tone and foley that track what is actually on screenambient, foley

Black Forest Labs single out facial expression, tying sound to physical events, and multilingual speech as the places their early evaluations came out strongest. Those are also the three that usually give AI video away, so they are the ones to test first on your own brief.

You set the frames. It fills the gaps.

Attached images are not mood references here. They are placed on screen pixel for pixel, and how many you attach changes what the model is being asked to do. One image opens the clip. Two set the first and last frame, so you are specifying a transition rather than hoping for one. Three to ten become a storyboard, spread evenly through the clip, and the model works out the motion between the beats you fixed. That is closer to directing than prompting.

the same woman, the same apple, four angles that agree with each otherconsistency
blocking a sequence out shot by shot before anything is finishedstoryboarding

A clip can also continue from the last frames of a previous one, so a sequence is chained rather than generated in one impossible take. Characters, setting and pacing carry across the join.

Draft first, then commit

Every prompt has a cheap version. Draft mode returns a fast preview at a fraction of the credits, meant for the four or five rounds where you are still working out whether the idea holds. When it does, re run the same prompt at full quality. The composition survives; the finish arrives.

koi between paper lanterns, draft passdraft
the same prompt, full qualityfull render

A range that does not read as AI video

The failure mode of most video models is that everything comes out looking like the same expensive commercial. FLUX 3 holds a look when you name one, and the useful looks are rarely cinematic. Camcorder footage, 1990s broadcast, claymation, anime, film noir, thermal camera. Naming the format does more work than any amount of asking for quality.

a 1990s news package, down to the tape softness and the mic tonebroadcast
claymation, with the fingerprints left instop motion
a ghost story by a campfire, carried mostly by the ambiencestorytelling

How to run it in Skiprr

Pick FLUX 3 in the model picker on a video production. Write the prompt in plain sentences rather than comma separated tags, because the prompt is read and expanded before generation, and say what the shot should sound like as well as what it should look like. Attach keyframes if you want to fix the beats, or feed the previous clip to continue a sequence. Clips run 5 to 20 seconds, at 720p or 1080p, in any of eight aspect ratios or on auto.

prompt shape

Handheld camcorder footage, slightly soft, of a corner bakery at six in the morning. The baker slides a tray of rolls onto the rack and wipes their hands on an apron. Warm overhead light, steam on the window, the street still dark outside. Audio: the scrape of the tray, an extractor fan running, rain on the awning, and the baker saying "first batch is out" to nobody in particular.

One thing to be straight about: Black Forest Labs call this an early access preview and describe their own evaluations as preliminary. So FLUX 3 leads the picker but Seedance 2.5 stays preselected as the default for now. If a client deliverable has to land today, start on the default. If you want the newest model in the world on it, FLUX 3 is one click away, and your brand brain still runs the show either way, so references, voice and guardrails carry into every generation.

see it running on your brand
// keep walking the studio
Back to the front door